Open WebUI on Thumper
Open WebUI is a self-hosted ChatGPT-style interface for local LLMs. It connects to Ollama for model management and provides a polished chat experience with conversation history, RAG upload, and multi-model switching.
Manual vs Thumper Install
Setting up Open WebUI manually means installing Ollama separately, pulling models from the CLI, and configuring the connection. Thumper handles the dependency chain for you.
| Step | Manual Install | Thumper Install | Time Saved |
|---|---|---|---|
| 1. Install Ollama | Download + install separately | Auto-started as dependency | ~3 min |
| 2. Install Open WebUI | pip install or docker | Click Install | ~2 min |
| 3. Download models | ollama pull | Auto-pull from pack | ~5 min |
| 4. Configure connection | Set OLLAMA_BASE_URL | Automatic | ~2 min |
| 5. Launch | open-webui serve | Click Launch | ~1 min |
Ollama Integration
Open WebUI relies on Ollama as its inference backend. Thumper treats Ollama as a managed dependency — it starts Ollama before launching Open WebUI, monitors its health, and auto-pulls any models declared in the app’s manifest.
# Key environment variables set by Thumper:OLLAMA_BASE_URL=http://127.0.0.1:11434WEBUI_AUTH=FalseDATA_DIR=~/.local/share/tr-desktop/apps/open-webui/data
| GPU | Status | Backend | Image Gen | LLM Speed | Notes |
|---|---|---|---|---|---|
| RTX 4090 (24 GB) | Full | CUDA 12.4 | ~3s | ~120 tok/s | Fastest consumer GPU |
| RTX 4070 (12 GB) | Full | CUDA 12.4 | ~8s | ~80 tok/s | Great balance of price/performance |
| RTX 3060 (12 GB) | Full | CUDA 11.8 | ~15s | ~45 tok/s | 12 GB VRAM at budget price |
| GTX 1660 (6 GB) | Partial | CUDA 11.8 | ~30s | ~20 tok/s | 6 GB limits model size |
| RX 7900 XT (20 GB) | Full | ROCm 6.2 | ~6s | ~90 tok/s | Best AMD option, large VRAM |
| RX 7600 (8 GB) | Full | ROCm 6.2 | ~18s | ~40 tok/s | Budget AMD with ROCm support |
| Radeon 780M APU (8 GB shared) | Partial | ROCm 6.2 | ~45s | ~15 tok/s | BF16 only, 5-min MIOpen warmup |
| Arc A770 (16 GB) | Partial | oneAPI/IPEX | ~20s | ~35 tok/s | Requires oneAPI runtime |
| M2 Pro (16 GB unified) | Full | MPS (Metal) | ~12s | ~50 tok/s | Unified memory, no discrete VRAM limit |
| M1 (8 GB unified) | Partial | MPS (Metal) | ~35s | ~25 tok/s | 8 GB tight for SDXL |
| M3 Max (36 GB unified) | Full | MPS (Metal) | ~8s | ~70 tok/s | Runs large models easily |
| CPU only (no GPU) | Partial | CPU fallback | ~180s | ~5 tok/s | Works but 10-50x slower |
Model Configuration
The default model pack includes 8 LLM options across 4 size tiers: 4B, 8B, 14B, and 32B parameters, each available in chat and code variants. The right choice depends on your available RAM and desired response speed.
Switching Models
Use the model dropdown in the Open WebUI chat interface to switch between any pulled model. You can change models mid-conversation — the chat history is preserved.
Adding Custom Models
Pull additional models from Ollama’s model library using the CLI, or declare them in the app’s Thumper manifest for automatic provisioning:
# Pull a model manually:ollama pull llama3.1:8b# Or add to the manifest for auto-pull:ollama_models:- llama3.1:8b- codellama:7b
Key Features
- Conversation history with search
- RAG: upload PDF/documents for context
- Multi-model switching in single conversation
- Markdown rendering with code highlighting
- System prompt customization
- OpenAI-compatible API at port 8080
GPU Tips
Open WebUI itself doesn’t use GPU — Ollama does the inference. GPU tips apply to the Ollama backend.
- NVIDIA — works out of the box with CUDA
- AMD — ROCm support via Ollama, set HSA_OVERRIDE_GFX_VERSION if needed
- CPU — works but slower; 4B models recommended for responsive chat
Key Takeaways
RAG & Document Upload
Open WebUI supports Retrieval-Augmented Generation (RAG) — upload documents and ask questions about them. The LLM retrieves relevant passages from your documents to ground its responses.
Supported Formats
- PDF (.pdf)
- Plain text (.txt, .md, .csv)
- Word documents (.docx)
- Web pages (paste URL)
Workflow
- Click the file upload icon in the chat input area
- Select one or more documents (max ~50 MB each)
- Wait for the embedding process to complete (a few seconds per document)
- Ask questions — the LLM will cite relevant passages from your documents
Backup & Export
Open WebUI stores all conversation data locally. You can export and back up your chats.
Chat Data Location
# All chat history and settings are stored here:~/.local/share/tr-desktop/apps/open-webui/data/# Key files:data/webui.db # SQLite database with all chatsdata/uploads/ # Uploaded RAG documents
Exporting Chats
Use the Open WebUI settings panel to export individual conversations or all chats as JSON. You can also back up the entire data directory for a full backup.
Custom Models
Beyond the default model pack, you can add custom models from multiple sources:
Private HuggingFace Models
If you have access to private or gated HuggingFace models, pull them through Ollama with your HF token:
# Set your HuggingFace token:OLLAMA_HF_TOKEN=hf_xxxxx ollama pull hf.co/your-org/your-model
Local GGUF Files
Load any GGUF model file directly into Ollama with a Modelfile:
# Create a Modelfile:echo 'FROM /path/to/your-model.gguf' > Modelfileollama create my-custom-model -f Modelfile# The model appears in Open WebUI's dropdown immediately
Troubleshooting
Common issues and how to fix them:
| Symptom | Cause | Fix |
|---|---|---|
| Port 11434 conflict | Another Ollama instance running | Kill existing process or let Thumper use it |
| Model not in dropdown | Model not pulled | Run ollama pull or restart app |
| Empty UI on first launch | Backend not ready | Wait for health check (30s timeout) |
| Large file upload hangs | File too large for RAG | Split into smaller chunks (<50 MB) |
| Query timeout | Model too large for RAM | Switch to smaller model (4B or 8B) |