LTX 2.3 Audio as Standalone Speech Model
Integrate LTX 2.3 into your TTS pipeline for zero‑shot expressive voice cloning and 13‑language support.
Add LTX 2.3 to your TTS workflow and test multi‑language output.
Summary
LTX 2.3 audio is a standalone speech model that offers zero‑shot expressive voice cloning and speech generation. It uses an 8‑step distilled architecture with Gemma 3 12B text encoding and supports stage directions via <action> tags. The model runs at 1.5× real‑time on an RTX 4090, fits within 16 GB VRAM, and supports 13 languages at 48 kHz stereo output. Additionally, it can generate matching environmental sounds alongside speech. The Hugging Face page https://huggingface.co/ScenemaAI/scenema‑audio hosts the model and documentation. The model is designed for expressive TTS applications and can be integrated into existing pipelines. Its zero‑shot capability eliminates the need for fine‑tuning on new voices.
Key changes
- Zero‑shot expressive voice cloning and speech generation
- 8‑step distilled with Gemma 3 12B text encoding
- Stage directions via <action> tags
- Runs 1.5× real‑time on RTX 4090, 16 GB VRAM
- 13 languages, 48 kHz stereo output
- Generates matching environmental sounds