llama.cpp on Thumper
llama.cpp is the C/C++ inference engine that powers local LLM execution across the Thumper ecosystem. It runs inside Ollama, inside Jan, and as Thumper’s embedded inference backend for the built-in agent. Understanding how it works helps you choose the right models and get the best performance from your hardware.
Overview
llama.cpp is a pure C/C++ implementation of LLM inference. It loads GGUF model files and runs them on CPU, GPU, or a mix of both. It supports dozens of model architectures (LLaMA, Mistral, Qwen, Phi, Gemma, and more) through a single unified runtime.
Thumper uses llama.cpp in three ways:
- Via Ollama — Ollama wraps llama.cpp behind an OpenAI-compatible API, serving models for Open WebUI, SillyTavern, and other API consumers
- Via Jan — Jan bundles llama.cpp as its native inference engine for standalone desktop chat
- Embedded — Thumper links llama.cpp directly via the Rust
llama-cpp-2crate for the built-in agent, prompt enhancement, and background tasks
The embedded path uses dynamic GPU backend loading (GGML_BACKEND_DL), which means GPU support is loaded at runtime rather than compiled in. This allows a single Thumper binary to work with CUDA, ROCm, Metal, Vulkan, or CPU — depending on what’s available on the system.
GGUF Format
GGUF (GPT-Generated Unified Format) is the standard file format for llama.cpp models. A single .gguf file contains everything needed for inference:
- Model weights — the actual parameters, typically quantized to reduce size
- Tokenizer — vocabulary and merges, embedded in the file
- Metadata — architecture type, context length, quantization scheme, author, license
- Chat template — Jinja2 template for formatting messages (system, user, assistant roles)
The key advantage of GGUF is portability. Unlike safetensors or PyTorch checkpoints that require a separate tokenizer config, GGUF is fully self-contained. You can copy a single file to another machine and it just works.
Quantization
Quantization reduces model size and memory usage by storing weights at lower precision. llama.cpp supports many quantization levels, each trading quality for speed and size.
| Quantization | Size (7B model) | Speed | Quality | Use Case |
|---|---|---|---|---|
| Q4_K_M | ~4.5 GB | Fast | Good | Best all-around choice for most users |
| Q5_K_M | ~5.1 GB | Moderate | Very good | Slight quality bump over Q4 with more memory |
| Q6_K | ~5.9 GB | Moderate | Excellent | Near-lossless, good for code and reasoning |
| Q8_0 | ~7.5 GB | Slower | Near-lossless | Maximum quality at 8-bit, requires more RAM |
| FP16 | ~14 GB | Slowest | Lossless | Full precision, only for benchmarking or fine-tuning |
GPU Backends
Thumper uses llama.cpp’s dynamic backend loading system (GGML_BACKEND_DL). Instead of compiling separate binaries for each GPU vendor, a single binary loads the appropriate GPU backend at runtime by calling ggml_backend_load_all().
Supported Backends
| Backend | Hardware | Notes |
|---|---|---|
| CUDA | NVIDIA GPUs | Best supported, fastest on NVIDIA hardware |
| ROCm / HIP | AMD GPUs and APUs | Requires ROCm runtime; TheRock nightly for newest chips |
| Metal | Apple Silicon | Excellent performance on M-series chips, unified memory |
| Vulkan | Cross-platform fallback | Works on most GPUs, slower than native backends |
| CPU | Any system | Always available, uses AVX/AVX2/AVX-512 when present |
The dynamic loading approach means Thumper works on any hardware out of the box. If CUDA libraries are found, the CUDA backend is loaded automatically. If not, it falls back to the next available option. No configuration needed.
Ollama vs llama.cpp
Ollama and llama.cpp are not competing tools — Ollama is built on llama.cpp. The question is whether to use Ollama’s server or llama.cpp directly.
| Aspect | Ollama | llama.cpp (Direct) |
|---|---|---|
| API | OpenAI-compatible REST API | C API or CLI |
| Multi-model | Hot-swaps models on request | One model per process |
| Memory overhead | Server process (~50 MB baseline) | Minimal (embedded in host process) |
| Model format | Ollama registry (pulls by tag) | Raw .gguf files |
| Best for | Multi-app serving, chat UIs, API consumers | Embedded inference, single-purpose tasks, agents |
Key Takeaways
Context Length
Context length determines how much text the model can process at once. Longer context uses more memory and slows inference.
| Model | Max Context | Memory Impact |
|---|---|---|
| Llama 3.1 8B | 128K tokens | +0.5 GB per 4K tokens (Q4_K_M) |
| Mistral 7B | 32K tokens | +0.4 GB per 4K tokens (Q4_K_M) |
| Phi-3 Mini 3.8B | 128K tokens | +0.3 GB per 4K tokens (Q4_K_M) |
| Qwen3 4B | 32K tokens | +0.3 GB per 4K tokens (Q4_K_M) |
| Gemma 2 9B | 8K tokens | +0.5 GB per 4K tokens (Q4_K_M) |
Sampling Parameters
Sampling controls how the model picks the next token. These parameters affect creativity, coherence, and response diversity.
Temperature
Controls randomness. Lower values (0.1–0.3) produce focused, deterministic output. Higher values (0.7–1.0) produce more creative, varied output. Default: 0.8.
Top-p (Nucleus Sampling)
Keeps only tokens whose cumulative probability exceeds the threshold. A top_p of 0.9 means the model considers the smallest set of tokens covering 90% probability. Default: 0.9.
Top-k
Limits the model to the k most likely tokens at each step. Top_k of 40 means only the 40 most probable tokens are considered. Lower values = more focused. Default: 40.
Recommended Presets
- Code generation — temperature=0.2, top_p=0.9, top_k=20
- Chat / conversation — temperature=0.7, top_p=0.9, top_k=40
- Creative writing — temperature=1.0, top_p=0.95, top_k=60
- Factual Q&A — temperature=0.1, top_p=0.8, top_k=10
Benchmark Methodology
When measuring llama.cpp performance, follow these practices for reproducible results:
- Use a fixed prompt (512+ tokens input) for consistency
- Run at least 3 iterations and report the median
- Report both prompt processing speed (tokens/s) and generation speed (tokens/s) separately
- Specify exact model, quantization, context length, GPU layers offloaded, and backend
- Close other GPU-consuming applications during benchmarks
- For AMD: ensure MIOpen Find-DB is warm (first run is always slower)