Briefing

Stable Audio 3 Release Offers Open‑Weight Models for Music and SFX

ai-dev
by Louisa Marshall ·

Download the open‑weight Small and Medium models from Hugging Face, test the new variable‑length generation API, and experiment with LoRa fine‑tuning for custom audio datasets.

What to do now

Download the Small and Medium weights from Hugging Face, test variable‑length generation, and experiment with LoRa fine‑tuning for your custom audio dataset.

Summary

Stability AI has launched Stable Audio 3, a new family of text‑to‑audio models that are fully open‑weight and free to use under the Stability AI Community License. The release includes three distinct models: Stable Audio 3 Small Music, which can generate up to two minutes of music; Stable Audio 3 Small SFX, also capped at two minutes but focused on sound‑effects; and Stable Audio 3 Medium, which extends the maximum length to six minutes and twenty seconds. The Medium model is engineered to run inference in seconds on NVIDIA GPUs, while the Small models are lightweight enough to operate efficiently on standard CPUs, making the technology accessible to a wide range of users.

The open‑weight approach is a deliberate move by Stability AI to encourage community experimentation and fine‑tuning. A GitHub repository has been published that contains inference code and tools for LoRA fine‑tuning, allowing developers to adapt the models to specific styles or use cases. Two academic papers accompany the release, detailing the underlying architecture and the SAME autoencoder that powers the audio generation. A public demo link is also provided, giving users a quick way to test the models before downloading.

The extended duration of the Medium model opens new creative possibilities for longer audio content, such as full‑length tracks or extended sound‑effect sequences, which were previously limited by shorter model capacities. Stability AI states that the models are ready for integration into production pipelines or creative projects, and they can be downloaded from Hugging Face or cloned directly from the GitHub repository. The company emphasizes that the models can be used for personal and creative projects without royalties, fostering a collaborative ecosystem around audio generation.

By making the weights public and providing comprehensive tooling, Stability AI aims to lower the barrier to entry for audio AI and stimulate innovation across the music, gaming, and media industries. The release is expected to accelerate experimentation and adoption of text‑to‑audio technology in both hobbyist and professional settings.

Key changes

  • Four new models: Small SFX, Small, Medium, Large
  • Variable‑length generation up to 6 min 20 s (Small up to 2 min, Medium & Large >6 min)
  • Full music composition on‑device via Small model
  • LoRa fine‑tuning documentation and weights for Small & Medium
  • Audio inpainting: single‑segment, multi‑segment, causal continuation
  • Open‑weight Small SFX/Small/Medium; Large via API/enterprise
  • Enterprise license for >$1 M revenue with commercial coverage and indemnification

Affects

internal

Source angles · 2 perspectives

Stability AI Blog
Independent angle

Meet Stable Audio 3.0, the model family built for artistic experimentation with open-weight models

Open
r/StableDiffusion
Independent angle

Stable Audio 3 Release: Open‑Weights Models for Music and SFX

Open

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting