Stable Audio 3 Release Offers Open‑Weight Models for Music and SFX
Download the open‑weight Small and Medium models from Hugging Face, test the new variable‑length generation API, and experiment with LoRa fine‑tuning for custom audio datasets.
Download the Small and Medium weights from Hugging Face, test variable‑length generation, and experiment with LoRa fine‑tuning for your custom audio dataset.
Summary
Stability AI has launched Stable Audio 3, a new family of text‑to‑audio models that are fully open‑weight and free to use under the Stability AI Community License. The release includes three distinct models: Stable Audio 3 Small Music, which can generate up to two minutes of music; Stable Audio 3 Small SFX, also capped at two minutes but focused on sound‑effects; and Stable Audio 3 Medium, which extends the maximum length to six minutes and twenty seconds. The Medium model is engineered to run inference in seconds on NVIDIA GPUs, while the Small models are lightweight enough to operate efficiently on standard CPUs, making the technology accessible to a wide range of users.
The open‑weight approach is a deliberate move by Stability AI to encourage community experimentation and fine‑tuning. A GitHub repository has been published that contains inference code and tools for LoRA fine‑tuning, allowing developers to adapt the models to specific styles or use cases. Two academic papers accompany the release, detailing the underlying architecture and the SAME autoencoder that powers the audio generation. A public demo link is also provided, giving users a quick way to test the models before downloading.
The extended duration of the Medium model opens new creative possibilities for longer audio content, such as full‑length tracks or extended sound‑effect sequences, which were previously limited by shorter model capacities. Stability AI states that the models are ready for integration into production pipelines or creative projects, and they can be downloaded from Hugging Face or cloned directly from the GitHub repository. The company emphasizes that the models can be used for personal and creative projects without royalties, fostering a collaborative ecosystem around audio generation.
By making the weights public and providing comprehensive tooling, Stability AI aims to lower the barrier to entry for audio AI and stimulate innovation across the music, gaming, and media industries. The release is expected to accelerate experimentation and adoption of text‑to‑audio technology in both hobbyist and professional settings.
Key changes
- Four new models: Small SFX, Small, Medium, Large
- Variable‑length generation up to 6 min 20 s (Small up to 2 min, Medium & Large >6 min)
- Full music composition on‑device via Small model
- LoRa fine‑tuning documentation and weights for Small & Medium
- Audio inpainting: single‑segment, multi‑segment, causal continuation
- Open‑weight Small SFX/Small/Medium; Large via API/enterprise
- Enterprise license for >$1 M revenue with commercial coverage and indemnification