ACE-Step on Thumper

ACE-Step generates music from text descriptions. Enter a prompt like "upbeat jazz piano" and get a full audio track. It runs on Gradio and requires Python 3.11 exactly.

Manual vs Thumper Install

ACE-Step has a complex dependency chain: Python 3.11 specifically, PyTorch with GPU support, ffmpeg for audio encoding, and multiple model downloads. Thumper handles every step.

StepManual InstallThumper InstallTime Saved
1. Clone repositorygit clone + cdClick Install~1 min
2. Create Python 3.11 venvInstall Python 3.11 + venvAutomatic (version pinned)~5 min
3. Install PyTorchpip install torch (GPU-specific)GPU auto-detected~5 min
4. Install dependenciespip install + audio libsAutomatic~3 min
5. Download modelsManual HF download (~8 GB)Model pack auto-download~10 min
6. Install ffmpegapt/brew/choco installSystem dep auto-checked~2 min
7. Configure GPU flagsResearch + set env varsPlatform patches applied~15 min

VRAM Tiers

ACE-Step’s quality scales with available VRAM. The language model component (used for prompt understanding) is the main differentiator between tiers.

GPUStatusBackendImage GenLLM SpeedNotes
RTX 4090 (24 GB)FullCUDA 12.4~3s~120 tok/sFastest consumer GPU
RTX 4070 (12 GB)FullCUDA 12.4~8s~80 tok/sGreat balance of price/performance
RTX 3060 (12 GB)FullCUDA 11.8~15s~45 tok/s12 GB VRAM at budget price
GTX 1660 (6 GB)PartialCUDA 11.8~30s~20 tok/s6 GB limits model size
RX 7900 XT (20 GB)FullROCm 6.2~6s~90 tok/sBest AMD option, large VRAM
RX 7600 (8 GB)FullROCm 6.2~18s~40 tok/sBudget AMD with ROCm support
Radeon 780M APU (8 GB shared)PartialROCm 6.2~45s~15 tok/sBF16 only, 5-min MIOpen warmup
Arc A770 (16 GB)PartialoneAPI/IPEX~20s~35 tok/sRequires oneAPI runtime
M2 Pro (16 GB unified)FullMPS (Metal)~12s~50 tok/sUnified memory, no discrete VRAM limit
M1 (8 GB unified)PartialMPS (Metal)~35s~25 tok/s8 GB tight for SDXL
M3 Max (36 GB unified)FullMPS (Metal)~8s~70 tok/sRuns large models easily
CPU only (no GPU)PartialCPU fallback~180s~5 tok/sWorks but 10-50x slower
TierVRAMLanguage ModelQuality
Tier 1≤6 GBDiT only (no language model)Basic
Tier 26–12 GB0.6B LM (1.2 GB)Good
Tier 312–16 GB1.7B LM (3.5 GB)Best balance
Tier 4≥16 GB4B LM (8.0 GB)Highest

Key Takeaways

Thumper auto-selects the appropriate tier based on detected VRAM. You can override this in the app settings if needed.

Model Configuration

ACE-Step uses a Python runtime with Gradio on port 7860. The health check timeout is 180 seconds (longer than most apps) because model loading is slow on first launch.

bash
# Runtime configuration:
runtime: python
port: 7860 # Gradio default
health_check_timeout: 180 # 3 min for model loading
python_version: "3.11" # Exact version required
# System dependency:
system_deps:
- ffmpeg # Required for audio encoding

GPU-Specific Tips

ACE-Step’s DiT model is sensitive to precision and memory management. GPU configuration matters more here than for most apps.

NVIDIA (CUDA)

  • Fast inference out of the box with CUDA 11.8+
  • FP16 works correctly on all NVIDIA GPUs
  • RTX 4090 generates a 10-second clip in ~2 seconds

AMD Discrete GPU (ROCm)

  • Requires ROCm 5.7+
  • MIOPEN_FIND_MODE=2 is critical — without it, MIOpen re-benchmarks every launch (~8 min warmup)
  • First run builds the MIOpen Find-DB (5–8 min); subsequent runs are fast

AMD APU (Integrated Graphics)

  • BF16 is mandatory — FP16 causes NaN in DiT attention due to value overflow past FP16 max (~65504)
  • CPU text encoder frees GPU memory for diffusion
  • Smart memory (default) keeps the diffusion model cached between stages
NEVER use --lowvram on AMD APU. FP16 causes NaN — Thumper auto-enforces BF16.

Apple Silicon (MPS)

  • MPS backend works for most operations
  • May need FP32 fallback for some attention ops
  • M2 Max generates a 10-second clip in ~8 seconds

CPU Only

  • Very slow — basic generation only
  • Short clips (10s) recommended to keep generation time manageable

Performance Benchmarks

Generation time scales with clip duration. These benchmarks use the default model configuration at each GPU’s optimal tier.

GPU10s Clip30s Clip60s Clip
RTX 4090~2s~5s~10s
RX 7900 XT~4s~9s~18s
AMD 780M (APU)~43s~95sN/A
M2 Max~8s~20sN/A

Troubleshooting

Common issues and how to fix them:

SymptomCauseFix
CUDA out of memoryLanguage model too large for VRAM tierDrop to a lower VRAM tier or use CPU text encoder
NaN audio output (AMD)FP16 overflow in DiT attentionApp Detail → Settings tab → Precision → select BF16 (Thumper auto-enforces this on new installs)
MIOpen warmup takes 5–8 minBuilding kernel cache for AMD GPUNormal on first run; set MIOPEN_FIND_MODE=2 for fast mode
Very slow on APUShared memory bottleneckUse short clips (10–30s); ensure BF16 + CPU text encoder
Smart memory OOM during VAEResidual model data not evictedUse CPU text encoder or enable --disable-smart-memory
Python version errorWrong Python version (not 3.11)ACE-Step requires Python 3.11 exactly; install via pyenv or system package
macOS MPS crashUnsupported MPS operationForce FP32 fallback; update to latest PyTorch
Missing ffmpegSystem dependency not installedInstall via apt/brew/choco; Thumper checks on launch
GPU clock stuck at 804 MHzAMD power management throttlingSet performance power profile or force high clocks via sysfs
UI freeze on ≤6 GB tierGradio blocked by long generationUse short clips (10s max); increase VRAM allocation in BIOS