Open WebUI on Thumper

Open WebUI is a self-hosted ChatGPT-style interface for local LLMs. It connects to Ollama for model management and provides a polished chat experience with conversation history, RAG upload, and multi-model switching.

Manual vs Thumper Install

Setting up Open WebUI manually means installing Ollama separately, pulling models from the CLI, and configuring the connection. Thumper handles the dependency chain for you.

StepManual InstallThumper InstallTime Saved
1. Install OllamaDownload + install separatelyAuto-started as dependency~3 min
2. Install Open WebUIpip install or dockerClick Install~2 min
3. Download modelsollama pullAuto-pull from pack~5 min
4. Configure connectionSet OLLAMA_BASE_URLAutomatic~2 min
5. Launchopen-webui serveClick Launch~1 min

Ollama Integration

Open WebUI relies on Ollama as its inference backend. Thumper treats Ollama as a managed dependency — it starts Ollama before launching Open WebUI, monitors its health, and auto-pulls any models declared in the app’s manifest.

bash
# Key environment variables set by Thumper:
OLLAMA_BASE_URL=http://127.0.0.1:11434
WEBUI_AUTH=False
DATA_DIR=~/.local/share/tr-desktop/apps/open-webui/data
Thumper manages Ollama’s lifecycle automatically. If you already have Ollama running, Thumper detects it and connects.
GPUStatusBackendImage GenLLM SpeedNotes
RTX 4090 (24 GB)FullCUDA 12.4~3s~120 tok/sFastest consumer GPU
RTX 4070 (12 GB)FullCUDA 12.4~8s~80 tok/sGreat balance of price/performance
RTX 3060 (12 GB)FullCUDA 11.8~15s~45 tok/s12 GB VRAM at budget price
GTX 1660 (6 GB)PartialCUDA 11.8~30s~20 tok/s6 GB limits model size
RX 7900 XT (20 GB)FullROCm 6.2~6s~90 tok/sBest AMD option, large VRAM
RX 7600 (8 GB)FullROCm 6.2~18s~40 tok/sBudget AMD with ROCm support
Radeon 780M APU (8 GB shared)PartialROCm 6.2~45s~15 tok/sBF16 only, 5-min MIOpen warmup
Arc A770 (16 GB)PartialoneAPI/IPEX~20s~35 tok/sRequires oneAPI runtime
M2 Pro (16 GB unified)FullMPS (Metal)~12s~50 tok/sUnified memory, no discrete VRAM limit
M1 (8 GB unified)PartialMPS (Metal)~35s~25 tok/s8 GB tight for SDXL
M3 Max (36 GB unified)FullMPS (Metal)~8s~70 tok/sRuns large models easily
CPU only (no GPU)PartialCPU fallback~180s~5 tok/sWorks but 10-50x slower

Model Configuration

The default model pack includes 8 LLM options across 4 size tiers: 4B, 8B, 14B, and 32B parameters, each available in chat and code variants. The right choice depends on your available RAM and desired response speed.

Switching Models

Use the model dropdown in the Open WebUI chat interface to switch between any pulled model. You can change models mid-conversation — the chat history is preserved.

Adding Custom Models

Pull additional models from Ollama’s model library using the CLI, or declare them in the app’s Thumper manifest for automatic provisioning:

bash
# Pull a model manually:
ollama pull llama3.1:8b
# Or add to the manifest for auto-pull:
ollama_models:
- llama3.1:8b
- codellama:7b

Key Features

  • Conversation history with search
  • RAG: upload PDF/documents for context
  • Multi-model switching in single conversation
  • Markdown rendering with code highlighting
  • System prompt customization
  • OpenAI-compatible API at port 8080

GPU Tips

Open WebUI itself doesn’t use GPU — Ollama does the inference. GPU tips apply to the Ollama backend.

  • NVIDIA — works out of the box with CUDA
  • AMD — ROCm support via Ollama, set HSA_OVERRIDE_GFX_VERSION if needed
  • CPU — works but slower; 4B models recommended for responsive chat

Key Takeaways

For AMD APUs, Ollama automatically uses GPU layers when ROCm is available. If you see slow inference, confirm that Ollama is actually offloading to GPU by checking its logs.

RAG & Document Upload

Open WebUI supports Retrieval-Augmented Generation (RAG) — upload documents and ask questions about them. The LLM retrieves relevant passages from your documents to ground its responses.

Supported Formats

  • PDF (.pdf)
  • Plain text (.txt, .md, .csv)
  • Word documents (.docx)
  • Web pages (paste URL)

Workflow

  1. Click the file upload icon in the chat input area
  2. Select one or more documents (max ~50 MB each)
  3. Wait for the embedding process to complete (a few seconds per document)
  4. Ask questions — the LLM will cite relevant passages from your documents
For best results, use specific questions about the document content. RAG works well with technical documentation, research papers, and structured data.

Backup & Export

Open WebUI stores all conversation data locally. You can export and back up your chats.

Chat Data Location

bash
# All chat history and settings are stored here:
~/.local/share/tr-desktop/apps/open-webui/data/
# Key files:
data/webui.db # SQLite database with all chats
data/uploads/ # Uploaded RAG documents

Exporting Chats

Use the Open WebUI settings panel to export individual conversations or all chats as JSON. You can also back up the entire data directory for a full backup.

Custom Models

Beyond the default model pack, you can add custom models from multiple sources:

Private HuggingFace Models

If you have access to private or gated HuggingFace models, pull them through Ollama with your HF token:

bash
# Set your HuggingFace token:
OLLAMA_HF_TOKEN=hf_xxxxx ollama pull hf.co/your-org/your-model

Local GGUF Files

Load any GGUF model file directly into Ollama with a Modelfile:

bash
# Create a Modelfile:
echo 'FROM /path/to/your-model.gguf' > Modelfile
ollama create my-custom-model -f Modelfile
# The model appears in Open WebUI's dropdown immediately
Any model available in Ollama is automatically visible in Open WebUI. No additional configuration needed.

Troubleshooting

Common issues and how to fix them:

SymptomCauseFix
Port 11434 conflictAnother Ollama instance runningKill existing process or let Thumper use it
Model not in dropdownModel not pulledRun ollama pull or restart app
Empty UI on first launchBackend not readyWait for health check (30s timeout)
Large file upload hangsFile too large for RAGSplit into smaller chunks (<50 MB)
Query timeoutModel too large for RAMSwitch to smaller model (4B or 8B)