ACE-Step on Thumper
ACE-Step generates music from text descriptions. Enter a prompt like "upbeat jazz piano" and get a full audio track. It runs on Gradio and requires Python 3.11 exactly.
Manual vs Thumper Install
ACE-Step has a complex dependency chain: Python 3.11 specifically, PyTorch with GPU support, ffmpeg for audio encoding, and multiple model downloads. Thumper handles every step.
| Step | Manual Install | Thumper Install | Time Saved |
|---|---|---|---|
| 1. Clone repository | git clone + cd | Click Install | ~1 min |
| 2. Create Python 3.11 venv | Install Python 3.11 + venv | Automatic (version pinned) | ~5 min |
| 3. Install PyTorch | pip install torch (GPU-specific) | GPU auto-detected | ~5 min |
| 4. Install dependencies | pip install + audio libs | Automatic | ~3 min |
| 5. Download models | Manual HF download (~8 GB) | Model pack auto-download | ~10 min |
| 6. Install ffmpeg | apt/brew/choco install | System dep auto-checked | ~2 min |
| 7. Configure GPU flags | Research + set env vars | Platform patches applied | ~15 min |
VRAM Tiers
ACE-Step’s quality scales with available VRAM. The language model component (used for prompt understanding) is the main differentiator between tiers.
| GPU | Status | Backend | Image Gen | LLM Speed | Notes |
|---|---|---|---|---|---|
| RTX 4090 (24 GB) | Full | CUDA 12.4 | ~3s | ~120 tok/s | Fastest consumer GPU |
| RTX 4070 (12 GB) | Full | CUDA 12.4 | ~8s | ~80 tok/s | Great balance of price/performance |
| RTX 3060 (12 GB) | Full | CUDA 11.8 | ~15s | ~45 tok/s | 12 GB VRAM at budget price |
| GTX 1660 (6 GB) | Partial | CUDA 11.8 | ~30s | ~20 tok/s | 6 GB limits model size |
| RX 7900 XT (20 GB) | Full | ROCm 6.2 | ~6s | ~90 tok/s | Best AMD option, large VRAM |
| RX 7600 (8 GB) | Full | ROCm 6.2 | ~18s | ~40 tok/s | Budget AMD with ROCm support |
| Radeon 780M APU (8 GB shared) | Partial | ROCm 6.2 | ~45s | ~15 tok/s | BF16 only, 5-min MIOpen warmup |
| Arc A770 (16 GB) | Partial | oneAPI/IPEX | ~20s | ~35 tok/s | Requires oneAPI runtime |
| M2 Pro (16 GB unified) | Full | MPS (Metal) | ~12s | ~50 tok/s | Unified memory, no discrete VRAM limit |
| M1 (8 GB unified) | Partial | MPS (Metal) | ~35s | ~25 tok/s | 8 GB tight for SDXL |
| M3 Max (36 GB unified) | Full | MPS (Metal) | ~8s | ~70 tok/s | Runs large models easily |
| CPU only (no GPU) | Partial | CPU fallback | ~180s | ~5 tok/s | Works but 10-50x slower |
| Tier | VRAM | Language Model | Quality |
|---|---|---|---|
| Tier 1 | ≤6 GB | DiT only (no language model) | Basic |
| Tier 2 | 6–12 GB | 0.6B LM (1.2 GB) | Good |
| Tier 3 | 12–16 GB | 1.7B LM (3.5 GB) | Best balance |
| Tier 4 | ≥16 GB | 4B LM (8.0 GB) | Highest |
Key Takeaways
Model Configuration
ACE-Step uses a Python runtime with Gradio on port 7860. The health check timeout is 180 seconds (longer than most apps) because model loading is slow on first launch.
# Runtime configuration:runtime: pythonport: 7860 # Gradio defaulthealth_check_timeout: 180 # 3 min for model loadingpython_version: "3.11" # Exact version required# System dependency:system_deps:- ffmpeg # Required for audio encoding
GPU-Specific Tips
ACE-Step’s DiT model is sensitive to precision and memory management. GPU configuration matters more here than for most apps.
NVIDIA (CUDA)
- Fast inference out of the box with CUDA 11.8+
- FP16 works correctly on all NVIDIA GPUs
- RTX 4090 generates a 10-second clip in ~2 seconds
AMD Discrete GPU (ROCm)
- Requires ROCm 5.7+
MIOPEN_FIND_MODE=2is critical — without it, MIOpen re-benchmarks every launch (~8 min warmup)- First run builds the MIOpen Find-DB (5–8 min); subsequent runs are fast
AMD APU (Integrated Graphics)
- BF16 is mandatory — FP16 causes NaN in DiT attention due to value overflow past FP16 max (~65504)
- CPU text encoder frees GPU memory for diffusion
- Smart memory (default) keeps the diffusion model cached between stages
Apple Silicon (MPS)
- MPS backend works for most operations
- May need FP32 fallback for some attention ops
- M2 Max generates a 10-second clip in ~8 seconds
CPU Only
- Very slow — basic generation only
- Short clips (10s) recommended to keep generation time manageable
Performance Benchmarks
Generation time scales with clip duration. These benchmarks use the default model configuration at each GPU’s optimal tier.
| GPU | 10s Clip | 30s Clip | 60s Clip |
|---|---|---|---|
| RTX 4090 | ~2s | ~5s | ~10s |
| RX 7900 XT | ~4s | ~9s | ~18s |
| AMD 780M (APU) | ~43s | ~95s | N/A |
| M2 Max | ~8s | ~20s | N/A |
Troubleshooting
Common issues and how to fix them:
| Symptom | Cause | Fix |
|---|---|---|
| CUDA out of memory | Language model too large for VRAM tier | Drop to a lower VRAM tier or use CPU text encoder |
| NaN audio output (AMD) | FP16 overflow in DiT attention | App Detail → Settings tab → Precision → select BF16 (Thumper auto-enforces this on new installs) |
| MIOpen warmup takes 5–8 min | Building kernel cache for AMD GPU | Normal on first run; set MIOPEN_FIND_MODE=2 for fast mode |
| Very slow on APU | Shared memory bottleneck | Use short clips (10–30s); ensure BF16 + CPU text encoder |
| Smart memory OOM during VAE | Residual model data not evicted | Use CPU text encoder or enable --disable-smart-memory |
| Python version error | Wrong Python version (not 3.11) | ACE-Step requires Python 3.11 exactly; install via pyenv or system package |
| macOS MPS crash | Unsupported MPS operation | Force FP32 fallback; update to latest PyTorch |
| Missing ffmpeg | System dependency not installed | Install via apt/brew/choco; Thumper checks on launch |
| GPU clock stuck at 804 MHz | AMD power management throttling | Set performance power profile or force high clocks via sysfs |
| UI freeze on ≤6 GB tier | Gradio blocked by long generation | Use short clips (10s max); increase VRAM allocation in BIOS |