Jan on Thumper
Jan is a native desktop chat app powered by llama.cpp. Unlike web-based tools, it runs as a standalone application — no browser, no server, no port configuration.
What is Jan?
Jan is a native application (AppImage on Linux, DMG on macOS, EXE on Windows) that bundles llama.cpp as its inference engine. It is not a web server — the UI runs natively in its own window, not in your browser.
This makes Jan fundamentally different from Open WebUI or SillyTavern. There’s no Ollama dependency, no port to configure, and no URL to visit. You launch it and start chatting.
Installation
Install Jan from the Thumper catalog. Thumper downloads the correct binary for your platform from Jan’s GitHub releases and handles setup automatically.
Linux
Jan ships as an AppImage. Thumper downloads it, sets execute permissions, and registers it in the app launcher. No additional dependencies required.
macOS
Jan ships as a DMG. Thumper mounts, copies the app bundle to the managed app directory, and handles Gatekeeper approval.
Windows
Jan ships as an EXE installer. Thumper runs the installer silently and tracks the installation path for launching.
Model Downloads
Jan uses GGUF-format models exclusively. These are quantized weight files optimized for llama.cpp inference. You can download models directly through Jan’s built-in model browser.
Quantization Levels
| Quantization | Size (8B model) | Quality | Speed |
|---|---|---|---|
| Q2_K | ~3.0 GB | Low | Fastest |
| Q4_K_M | ~4.5 GB | Good | Fast |
| Q6_K | ~5.5 GB | Very good | Moderate |
| Q8_0 | ~7.5 GB | Near-lossless | Slow |
GPU Acceleration
llama.cpp auto-detects GPU capability (CUDA, ROCm, Metal) and offloads layers to GPU when available. You can control how many layers are offloaded via Jan’s settings.
# Environment variables for GPU control:LLAMA_N_GPU_LAYERS=99 # Offload all layers (default)LLAMA_N_GPU_LAYERS=0 # CPU onlyLLAMA_N_GPU_LAYERS=20 # Partial offload
- NVIDIA — CUDA auto-detected, full GPU offload by default
- AMD — ROCm support, may need HSA_OVERRIDE_GFX_VERSION for newer GPUs
- Apple Silicon — Metal backend, excellent performance on M-series chips
- CPU — works on any system, 4B models recommended for usable speed
Offline Usage
Jan is fully offline after model download. No account is required, no telemetry is sent, and no network connection is needed for inference. This makes it ideal for air-gapped environments or privacy-sensitive workflows.
Key Takeaways
Why GGUF?
Jan uses GGUF exclusively. Here is how it compares to other model formats:
| Format | Engine | Quantized | Self-Contained | Best For |
|---|---|---|---|---|
| GGUF | llama.cpp | Yes | Yes (includes tokenizer) | CPU/GPU inference, desktop apps |
| safetensors | PyTorch / HF Transformers | No (FP16/BF16) | No (needs config files) | Fine-tuning, Python pipelines |
| ONNX | ONNX Runtime | Optional | No (needs config) | NPU, cross-framework deployment |
Where to Find GGUF Models
- HuggingFace — Search for 'GGUF' on huggingface.co. TheBloke and bartowski provide pre-quantized versions of popular models.
- Ollama Library — ollama.com/library lists models by name and tag. Each tag maps to a specific GGUF quantization.
- Jan Model Hub — Jan’s built-in model browser includes curated GGUF models from HuggingFace.
Memory Calculator
Rough formula for estimating memory usage:
Memory (GB) = Model Size (GB) + Context MemoryContext Memory = (context_length * 2 * num_layers * head_dim) / 1e9# Simplified rule of thumb:Memory ~= GGUF file size + 0.5 GB (for 4K context)Memory ~= GGUF file size + 2.0 GB (for 32K context)
Examples
| Model | Quant | File Size | 4K Context | 32K Context |
|---|---|---|---|---|
| Llama 3.1 8B | Q4_K_M | 4.5 GB | ~5 GB | ~6.5 GB |
| Qwen3 4B | Q4_K_M | 2.4 GB | ~3 GB | ~4.5 GB |
| Mistral 7B | Q6_K | 5.9 GB | ~6.5 GB | ~8 GB |
Conversation Backup
Jan stores all conversation data in a local directory. To back up or migrate your threads:
# Jan thread storage location:~/.local/share/jan/threads/# Each thread is a directory with messages.jsonl:threads/thread_abc123/messages.jsonl # One JSON object per messagethread.json # Thread metadata (model, title)
To back up: copy the entire threads/ directory. To restore: paste it back and restart Jan.
Troubleshooting
Common issues and how to fix them:
| Symptom | Cause | Fix |
|---|---|---|
| AppImage won’t launch (Linux) | Missing execute permission | Run chmod +x Jan.AppImage |
| GPU not detected | Missing CUDA/ROCm drivers | Install GPU drivers; check Jan settings for GPU layer count |
| Model download hangs | Network timeout or CDN issue | Retry download; check firewall settings |
| CPU inference very slow | No GPU offload; large model | Use a smaller model (Q4_K_M at 4B) or enable GPU layers |
| Crash on macOS ARM | Rosetta conflict or code signing | Download the native ARM build; clear quarantine attribute |
| Chat history lost | Data directory moved or corrupted | Check Jan data directory in settings; restore from backup |
| OOM with 8 GB RAM | Model + OS exceeds available memory | Use Q2_K or Q4_K_M quantization; close other apps |