GPU Detection

Stable

5-minute read

Your SystemUnknown OS • GPU not detected

In Plain English

When you launch an AI app, Thumper-Run automatically figures out what GPU you have and configures the app to use it. You don’t need to install CUDA drivers, set environment variables, or edit config files — it just works.

Think of it like plugging in a printer: your OS detects the hardware and loads the right driver. Thumper-Run does the same thing for GPUs, but for AI workloads.

How It Works

GPU detection happens in four steps, all before the app’s first process starts.

Step 1: Hardware Scan

Thumper-Run queries the system for available GPUs using platform-native APIs: NVML for NVIDIA, ROCm SMI for AMD, Metal for Apple, and Level Zero for Intel.

Step 2: Backend Selection

Based on the detected GPU, Thumper-Run picks the correct compute backend: CUDA for NVIDIA, ROCm for AMD, MPS for Apple Silicon, or CPU fallback.

Step 3: Patch Application

If the app manifest has GPU-conditional patches (e.g., ROCm config overrides), they are applied automatically before launch.

Step 4: Environment Injection

GPU-specific environment variables are injected into the app’s process (e.g., CUDA_VISIBLE_DEVICES, HSA_OVERRIDE_GFX_VERSION).

Detection Output

json
{
"vendor": "AMD",
"name": "Radeon RX 7900 XT",
"vram_mb": 20480,
"backend": "rocm",
"backend_version": "6.2",
"compute_capability": "gfx1100",
"driver_version": "24.10.3"
}

VRAM Guide

Different AI tasks require different amounts of GPU memory (VRAM). This table shows minimum requirements for common workloads.

TaskMin VRAMRecommendedNotes
Chat (LLM)CPU / 0 GB4 GBQuantized 8B models run on CPU; GPU speeds up 5–10x
Image Generation4 GB6 GBSDXL needs ~6 GB; SD 1.5 fits in 4 GB
Music Generation2 GB4 GBACE-Step works with 2 GB; 4 GB for longer songs
Video Generation8 GB12 GBLTX Video needs 8 GB min; longer clips need more
AMD APU Users: Integrated GPUs (e.g., Radeon 780M) share system RAM. They work, but you must use BF16 precision — FP16 causes NaN errors in some models. Expect a 5–10 minute MIOpen warmup on first run.

How much VRAM do you need? →

Why It Matters

  • Zero Configuration — No manual driver setup, CUDA version matching, or environment variable juggling.
  • Cross-Platform — The same app manifest works on NVIDIA, AMD, Apple Silicon, and CPU-only machines.
  • Optimal Performance — GPU-specific patches tune settings for your exact hardware.
  • Graceful Fallback — If no GPU is found, apps fall back to CPU automatically (slower but functional).
  • User Feedback — The UI shows detected GPU info so users know what’s happening.

Deep Dive

For implementation details including the detection code paths, VRAM-aware model variant selection, and the full GPU compatibility matrix, see the Platform Architecture page.

App developers can control GPU behavior through manifest fields documented in the Runtime Configuration section. Use gpu_required: true to warn users, and conditional patches to swap backends per accelerator.

Key Takeaways

  • GPU detection is automatic — no manual setup needed
  • Supports NVIDIA (CUDA), AMD (ROCm), Apple (MPS), Intel (oneAPI), and CPU fallback
  • Detection happens before app launch, not at runtime
  • Manifest patches can customize behavior per GPU vendor
  • VRAM requirements vary by task: chat is cheapest, video is most demanding
  • AMD APU users must use BF16 — FP16 causes NaN errors