Jan on Thumper

Jan is a native desktop chat app powered by llama.cpp. Unlike web-based tools, it runs as a standalone application — no browser, no server, no port configuration.

What is Jan?

Jan is a native application (AppImage on Linux, DMG on macOS, EXE on Windows) that bundles llama.cpp as its inference engine. It is not a web server — the UI runs natively in its own window, not in your browser.

This makes Jan fundamentally different from Open WebUI or SillyTavern. There’s no Ollama dependency, no port to configure, and no URL to visit. You launch it and start chatting.

Installation

Install Jan from the Thumper catalog. Thumper downloads the correct binary for your platform from Jan’s GitHub releases and handles setup automatically.

Linux

Jan ships as an AppImage. Thumper downloads it, sets execute permissions, and registers it in the app launcher. No additional dependencies required.

macOS

Jan ships as a DMG. Thumper mounts, copies the app bundle to the managed app directory, and handles Gatekeeper approval.

Windows

Jan ships as an EXE installer. Thumper runs the installer silently and tracks the installation path for launching.

Model Downloads

Jan uses GGUF-format models exclusively. These are quantized weight files optimized for llama.cpp inference. You can download models directly through Jan’s built-in model browser.

Q4_K_M is the sweet spot — ~4.5 GB for an 8B model, good quality, fast inference.

Quantization Levels

QuantizationSize (8B model)QualitySpeed
Q2_K~3.0 GBLowFastest
Q4_K_M~4.5 GBGoodFast
Q6_K~5.5 GBVery goodModerate
Q8_0~7.5 GBNear-losslessSlow

GPU Acceleration

llama.cpp auto-detects GPU capability (CUDA, ROCm, Metal) and offloads layers to GPU when available. You can control how many layers are offloaded via Jan’s settings.

bash
# Environment variables for GPU control:
LLAMA_N_GPU_LAYERS=99 # Offload all layers (default)
LLAMA_N_GPU_LAYERS=0 # CPU only
LLAMA_N_GPU_LAYERS=20 # Partial offload
  • NVIDIA — CUDA auto-detected, full GPU offload by default
  • AMD — ROCm support, may need HSA_OVERRIDE_GFX_VERSION for newer GPUs
  • Apple Silicon — Metal backend, excellent performance on M-series chips
  • CPU — works on any system, 4B models recommended for usable speed

Offline Usage

Jan is fully offline after model download. No account is required, no telemetry is sent, and no network connection is needed for inference. This makes it ideal for air-gapped environments or privacy-sensitive workflows.

Key Takeaways

Once a GGUF model is downloaded, Jan never needs internet access again. All inference runs locally on your hardware.

Why GGUF?

Jan uses GGUF exclusively. Here is how it compares to other model formats:

FormatEngineQuantizedSelf-ContainedBest For
GGUFllama.cppYesYes (includes tokenizer)CPU/GPU inference, desktop apps
safetensorsPyTorch / HF TransformersNo (FP16/BF16)No (needs config files)Fine-tuning, Python pipelines
ONNXONNX RuntimeOptionalNo (needs config)NPU, cross-framework deployment
GGUF is the only format that bundles weights, tokenizer, and chat template in a single file. Copy one file, run on any machine.

Where to Find GGUF Models

  • HuggingFace — Search for 'GGUF' on huggingface.co. TheBloke and bartowski provide pre-quantized versions of popular models.
  • Ollama Library — ollama.com/library lists models by name and tag. Each tag maps to a specific GGUF quantization.
  • Jan Model Hub — Jan’s built-in model browser includes curated GGUF models from HuggingFace.

Memory Calculator

Rough formula for estimating memory usage:

Memory (GB) = Model Size (GB) + Context Memory
Context Memory = (context_length * 2 * num_layers * head_dim) / 1e9
# Simplified rule of thumb:
Memory ~= GGUF file size + 0.5 GB (for 4K context)
Memory ~= GGUF file size + 2.0 GB (for 32K context)

Examples

ModelQuantFile Size4K Context32K Context
Llama 3.1 8BQ4_K_M4.5 GB~5 GB~6.5 GB
Qwen3 4BQ4_K_M2.4 GB~3 GB~4.5 GB
Mistral 7BQ6_K5.9 GB~6.5 GB~8 GB
If memory exceeds your RAM, reduce context length in Jan settings or use a smaller quantization (Q4_K_M instead of Q6_K).

Conversation Backup

Jan stores all conversation data in a local directory. To back up or migrate your threads:

bash
# Jan thread storage location:
~/.local/share/jan/threads/
# Each thread is a directory with messages.jsonl:
threads/
thread_abc123/
messages.jsonl # One JSON object per message
thread.json # Thread metadata (model, title)

To back up: copy the entire threads/ directory. To restore: paste it back and restart Jan.

Troubleshooting

Common issues and how to fix them:

SymptomCauseFix
AppImage won’t launch (Linux)Missing execute permissionRun chmod +x Jan.AppImage
GPU not detectedMissing CUDA/ROCm driversInstall GPU drivers; check Jan settings for GPU layer count
Model download hangsNetwork timeout or CDN issueRetry download; check firewall settings
CPU inference very slowNo GPU offload; large modelUse a smaller model (Q4_K_M at 4B) or enable GPU layers
Crash on macOS ARMRosetta conflict or code signingDownload the native ARM build; clear quarantine attribute
Chat history lostData directory moved or corruptedCheck Jan data directory in settings; restore from backup
OOM with 8 GB RAMModel + OS exceeds available memoryUse Q2_K or Q4_K_M quantization; close other apps