audio.cpp 0.3 Release Adds Supertonic 3 and New TTS Models
Bump audio.cpp to 0.3 to unlock Supertonic 3 and other new TTS models with 200×+ CUDA real‑time speed.
Bump audio.cpp to 0.3 to unlock Supertonic 3 and other new TTS models with 200×+ CUDA real‑time speed.
Summary
audio.cpp has released version 0.3, adding five new TTS models: Supertonic 3, MOSS‑TTS‑Local, MOSS‑TTS‑Nano, IndexTTS2, and Irodori‑TTS. Supertonic 3 achieves over 200× real‑time speed on CUDA RTX 5090, 6× on CPU, and a 47 ms TTFT in CUDA streaming mode, generating 10 hours of audio in roughly three minutes on an RTX 5090. The release uses an ONNX implementation, requiring reverse‑engineering of the inference path around safetensors weights. GGUF support is added and will roll out model‑by‑model. The C++ implementation outperforms the Python counterpart by up to 5.65× on long‑form tests, and the overall performance gains are attributed to efficient GPU utilization. The project remains open source, encouraging contributions to model coverage, streaming, backend compatibility, and UI layers.
Key changes
- Release 0.3 adds Supertonic 3, MOSS‑TTS‑Local, MOSS‑TTS‑Nano, IndexTTS2, Irodori‑TTS
- Supertonic 3 achieves 200×+ real‑time on CUDA RTX 5090
- Supertonic 3 achieves 6×+ real‑time on CPU
- Supertonic 3 TTFT 47 ms in CUDA streaming mode
- ONNX implementation requires reverse‑engineering of safetensors weights
- GGUF support added and will roll out model‑by‑model
- C++ outperforms Python by up to 5.65× on long‑form tests