llama.cpp on Thumper

llama.cpp is the C/C++ inference engine that powers local LLM execution across the Thumper ecosystem. It runs inside Ollama, inside Jan, and as Thumper’s embedded inference backend for the built-in agent. Understanding how it works helps you choose the right models and get the best performance from your hardware.

Overview

llama.cpp is a pure C/C++ implementation of LLM inference. It loads GGUF model files and runs them on CPU, GPU, or a mix of both. It supports dozens of model architectures (LLaMA, Mistral, Qwen, Phi, Gemma, and more) through a single unified runtime.

Thumper uses llama.cpp in three ways:

  • Via Ollama — Ollama wraps llama.cpp behind an OpenAI-compatible API, serving models for Open WebUI, SillyTavern, and other API consumers
  • Via Jan — Jan bundles llama.cpp as its native inference engine for standalone desktop chat
  • Embedded — Thumper links llama.cpp directly via the Rust llama-cpp-2 crate for the built-in agent, prompt enhancement, and background tasks

The embedded path uses dynamic GPU backend loading (GGML_BACKEND_DL), which means GPU support is loaded at runtime rather than compiled in. This allows a single Thumper binary to work with CUDA, ROCm, Metal, Vulkan, or CPU — depending on what’s available on the system.

GGUF Format

GGUF (GPT-Generated Unified Format) is the standard file format for llama.cpp models. A single .gguf file contains everything needed for inference:

  • Model weights — the actual parameters, typically quantized to reduce size
  • Tokenizer — vocabulary and merges, embedded in the file
  • Metadata — architecture type, context length, quantization scheme, author, license
  • Chat template — Jinja2 template for formatting messages (system, user, assistant roles)

The key advantage of GGUF is portability. Unlike safetensors or PyTorch checkpoints that require a separate tokenizer config, GGUF is fully self-contained. You can copy a single file to another machine and it just works.

GGUF files from HuggingFace work with both Ollama and Jan. You can download a .gguf file once and use it in multiple apps.

Quantization

Quantization reduces model size and memory usage by storing weights at lower precision. llama.cpp supports many quantization levels, each trading quality for speed and size.

QuantizationSize (7B model)SpeedQualityUse Case
Q4_K_M~4.5 GBFastGoodBest all-around choice for most users
Q5_K_M~5.1 GBModerateVery goodSlight quality bump over Q4 with more memory
Q6_K~5.9 GBModerateExcellentNear-lossless, good for code and reasoning
Q8_0~7.5 GBSlowerNear-losslessMaximum quality at 8-bit, requires more RAM
FP16~14 GBSlowestLosslessFull precision, only for benchmarking or fine-tuning
Q4_K_M is the sweet spot for most users — good quality at ~4.5 GB for a 7B model. Start here and only go higher if you notice quality issues for your specific task.

GPU Backends

Thumper uses llama.cpp’s dynamic backend loading system (GGML_BACKEND_DL). Instead of compiling separate binaries for each GPU vendor, a single binary loads the appropriate GPU backend at runtime by calling ggml_backend_load_all().

Supported Backends

BackendHardwareNotes
CUDANVIDIA GPUsBest supported, fastest on NVIDIA hardware
ROCm / HIPAMD GPUs and APUsRequires ROCm runtime; TheRock nightly for newest chips
MetalApple SiliconExcellent performance on M-series chips, unified memory
VulkanCross-platform fallbackWorks on most GPUs, slower than native backends
CPUAny systemAlways available, uses AVX/AVX2/AVX-512 when present

The dynamic loading approach means Thumper works on any hardware out of the box. If CUDA libraries are found, the CUDA backend is loaded automatically. If not, it falls back to the next available option. No configuration needed.

Ollama vs llama.cpp

Ollama and llama.cpp are not competing tools — Ollama is built on llama.cpp. The question is whether to use Ollama’s server or llama.cpp directly.

AspectOllamallama.cpp (Direct)
APIOpenAI-compatible REST APIC API or CLI
Multi-modelHot-swaps models on requestOne model per process
Memory overheadServer process (~50 MB baseline)Minimal (embedded in host process)
Model formatOllama registry (pulls by tag)Raw .gguf files
Best forMulti-app serving, chat UIs, API consumersEmbedded inference, single-purpose tasks, agents
Most users should use Ollama. It handles model management, concurrent requests, and provides an easy API. llama.cpp direct usage is for advanced cases: embedded inference in the Thumper agent, custom prompt enhancement nodes, or when you need to minimize overhead for a single dedicated model.

Key Takeaways

llama.cpp is the engine; Ollama is the service layer. Thumper uses both — Ollama for multi-app model serving, llama.cpp directly for embedded inference in the built-in agent and prompt enhancement.

Context Length

Context length determines how much text the model can process at once. Longer context uses more memory and slows inference.

ModelMax ContextMemory Impact
Llama 3.1 8B128K tokens+0.5 GB per 4K tokens (Q4_K_M)
Mistral 7B32K tokens+0.4 GB per 4K tokens (Q4_K_M)
Phi-3 Mini 3.8B128K tokens+0.3 GB per 4K tokens (Q4_K_M)
Qwen3 4B32K tokens+0.3 GB per 4K tokens (Q4_K_M)
Gemma 2 9B8K tokens+0.5 GB per 4K tokens (Q4_K_M)
Start with 4K context for chat tasks. Increase only when you need to process long documents. Most conversations fit comfortably in 4K tokens.

Sampling Parameters

Sampling controls how the model picks the next token. These parameters affect creativity, coherence, and response diversity.

Temperature

Controls randomness. Lower values (0.1–0.3) produce focused, deterministic output. Higher values (0.7–1.0) produce more creative, varied output. Default: 0.8.

Top-p (Nucleus Sampling)

Keeps only tokens whose cumulative probability exceeds the threshold. A top_p of 0.9 means the model considers the smallest set of tokens covering 90% probability. Default: 0.9.

Top-k

Limits the model to the k most likely tokens at each step. Top_k of 40 means only the 40 most probable tokens are considered. Lower values = more focused. Default: 40.

Recommended Presets

  • Code generation — temperature=0.2, top_p=0.9, top_k=20
  • Chat / conversation — temperature=0.7, top_p=0.9, top_k=40
  • Creative writing — temperature=1.0, top_p=0.95, top_k=60
  • Factual Q&A — temperature=0.1, top_p=0.8, top_k=10

Benchmark Methodology

When measuring llama.cpp performance, follow these practices for reproducible results:

  • Use a fixed prompt (512+ tokens input) for consistency
  • Run at least 3 iterations and report the median
  • Report both prompt processing speed (tokens/s) and generation speed (tokens/s) separately
  • Specify exact model, quantization, context length, GPU layers offloaded, and backend
  • Close other GPU-consuming applications during benchmarks
  • For AMD: ensure MIOpen Find-DB is warm (first run is always slower)
Thumper’s built-in benchmark scripts are in scripts/ltx2/benchmark.py. These follow the methodology above and produce structured JSON output for comparison.