GPU Detection
Stable5-minute read
In Plain English
When you launch an AI app, Thumper-Run automatically figures out what GPU you have and configures the app to use it. You don’t need to install CUDA drivers, set environment variables, or edit config files — it just works.
Think of it like plugging in a printer: your OS detects the hardware and loads the right driver. Thumper-Run does the same thing for GPUs, but for AI workloads.
How It Works
GPU detection happens in four steps, all before the app’s first process starts.
Step 1: Hardware Scan
Thumper-Run queries the system for available GPUs using platform-native APIs: NVML for NVIDIA, ROCm SMI for AMD, Metal for Apple, and Level Zero for Intel.
Step 2: Backend Selection
Based on the detected GPU, Thumper-Run picks the correct compute backend: CUDA for NVIDIA, ROCm for AMD, MPS for Apple Silicon, or CPU fallback.
Step 3: Patch Application
If the app manifest has GPU-conditional patches (e.g., ROCm config overrides), they are applied automatically before launch.
Step 4: Environment Injection
GPU-specific environment variables are injected into the app’s process (e.g., CUDA_VISIBLE_DEVICES, HSA_OVERRIDE_GFX_VERSION).
Detection Output
{"vendor": "AMD","name": "Radeon RX 7900 XT","vram_mb": 20480,"backend": "rocm","backend_version": "6.2","compute_capability": "gfx1100","driver_version": "24.10.3"}
VRAM Guide
Different AI tasks require different amounts of GPU memory (VRAM). This table shows minimum requirements for common workloads.
| Task | Min VRAM | Recommended | Notes |
|---|---|---|---|
| Chat (LLM) | CPU / 0 GB | 4 GB | Quantized 8B models run on CPU; GPU speeds up 5–10x |
| Image Generation | 4 GB | 6 GB | SDXL needs ~6 GB; SD 1.5 fits in 4 GB |
| Music Generation | 2 GB | 4 GB | ACE-Step works with 2 GB; 4 GB for longer songs |
| Video Generation | 8 GB | 12 GB | LTX Video needs 8 GB min; longer clips need more |
Why It Matters
- Zero Configuration — No manual driver setup, CUDA version matching, or environment variable juggling.
- Cross-Platform — The same app manifest works on NVIDIA, AMD, Apple Silicon, and CPU-only machines.
- Optimal Performance — GPU-specific patches tune settings for your exact hardware.
- Graceful Fallback — If no GPU is found, apps fall back to CPU automatically (slower but functional).
- User Feedback — The UI shows detected GPU info so users know what’s happening.
Deep Dive
For implementation details including the detection code paths, VRAM-aware model variant selection, and the full GPU compatibility matrix, see the Platform Architecture page.
App developers can control GPU behavior through manifest fields documented in the Runtime Configuration section. Use gpu_required: true to warn users, and conditional patches to swap backends per accelerator.
Key Takeaways
- GPU detection is automatic — no manual setup needed
- Supports NVIDIA (CUDA), AMD (ROCm), Apple (MPS), Intel (oneAPI), and CPU fallback
- Detection happens before app launch, not at runtime
- Manifest patches can customize behavior per GPU vendor
- VRAM requirements vary by task: chat is cheapest, video is most demanding
- AMD APU users must use BF16 — FP16 causes NaN errors