Briefing

LTX 2.3 Audio as Standalone Speech Model

ai-dev
by /u/Famous-Sport7862 ·

Integrate LTX 2.3 into your TTS pipeline for zero‑shot expressive voice cloning and 13‑language support.

What to do now

Add LTX 2.3 to your TTS workflow and test multi‑language output.

Summary

LTX 2.3 audio is a standalone speech model that offers zero‑shot expressive voice cloning and speech generation. It uses an 8‑step distilled architecture with Gemma 3 12B text encoding and supports stage directions via <action> tags. The model runs at 1.5× real‑time on an RTX 4090, fits within 16 GB VRAM, and supports 13 languages at 48 kHz stereo output. Additionally, it can generate matching environmental sounds alongside speech. The Hugging Face page https://huggingface.co/ScenemaAI/scenema‑audio hosts the model and documentation. The model is designed for expressive TTS applications and can be integrated into existing pipelines. Its zero‑shot capability eliminates the need for fine‑tuning on new voices.

Key changes

  • Zero‑shot expressive voice cloning and speech generation
  • 8‑step distilled with Gemma 3 12B text encoding
  • Stage directions via <action> tags
  • Runs 1.5× real‑time on RTX 4090, 16 GB VRAM
  • 13 languages, 48 kHz stereo output
  • Generates matching environmental sounds

Affects

none

Source angles · 2 perspectives

Black Forest Labs (Reddit)
Independent angle

LTX 2.3 audio as standalone speech model.

Open
r/StableDiffusion
Independent angle

LTX 2.3 Audio as Standalone Speech Model

Open

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting