Stable Audio 3.0: six-minute instrumental tracks from models trained only on licensed audio
Stable Audio 3.0 is a family of four audio generation models from Stability AI, trained entirely on licensed data: you describe a track or a sound effect in plain text, it returns a 44.1 kHz stereo file. Three of the four variants download free from Hugging Face, weights included. Tracks now run past six minutes, twice the ceiling of Stable Audio 2.0. The heavy lifting happens in the prompt, so descriptive writing pays off here.
- Three open-weight models, free to download
- Tracks past six minutes, length set per second
- Training data fully licensed, outputs owned by you
- Full offline composition, even on a phone
- No vocals or lyrics generated
- Large model gated behind the API and enterprise deals
- Local setup asks for some command-line comfort
Four models, from pocket sound effects to six-minute tracks
The Stable Audio 3.0 line-up counts four variants: Small SFX, Small, Medium and Large, all built on a latent diffusion model that renders 44.1 kHz stereo audio. Length is set per second: type "warm ambient pad, 90 BPM, six minutes" and the track comes out at that exact duration.
Audio inpainting is part of the package: keep the intro of a piece and regenerate the ending. And Small composes a full two-minute track on a phone, offline, where the older Stable Audio Open Small topped out at eleven seconds.
| Model | Maximum length | Access |
|---|---|---|
| 3.0 Small SFX | Short sound effects | Open weights (Hugging Face) |
| 3.0 Small | 2 minutes | Open weights, runs on mobile |
| 3.0 Medium | Over 6 minutes | Open weights (Hugging Face) |
| 3.0 Large | Over 6 minutes | Stability AI API, enterprise self-hosting |
Stable Audio 3.0 bets on fully licensed training data
Every Stable Audio 3.0 model was trained on fully licensed data, a rare stance in generated music. Under the Community License, your outputs belong to you, sale and distribution included; past one million dollars in annual revenue, the Enterprise license takes over and adds legal indemnification.
Here is where it gets interesting: AI music generators have split into two schools, the finished vocal song in the style of Suno or Udio, and the open instrumental engine you control end to end. Stable Audio 3.0 picks the second lane, model weights included, and both approaches sit comfortably in the same project.
Trying Stable Audio 3.0 without a credit card
The fastest test lives at stableaudio.com, where the free plan includes 6 generations per month on the current model. A Hugging Face demo space also works without an account, within the platform's quotas. Tinkerers will head straight for ComfyUI, which ships ready-made workflows (the weights are hefty, plan the download).
Locally, Small runs without a GPU at all through CoreML, and Medium makes do with roughly 5 GB of video memory using chunked decoding. Large stays server-side, behind Stability AI's paid API. Quotas and terms shift between releases, so give the official page a quick look before committing.
Frequently asked questions
Is Stable Audio 3.0 free?
Largely, yes: Small SFX, Small and Medium download free from Hugging Face and run locally with no subscription. The hosted site at stableaudio.com adds 6 free generations per month. Only Large, the strongest of the family, sits behind Stability AI's paid API or an enterprise agreement.
Can you sell music made with Stable Audio 3.0?
Under the Community License, your outputs are yours to distribute and sell, royalty-free. Organizations above one million dollars in annual revenue must move to the Enterprise license, which comes with legal indemnification. On the hosted site, check the rights attached to your plan before publishing anything.
Does Stable Audio 3.0 generate vocals?
No. The family produces instrumental music, ambient textures and sound effects, but no singing and no lyrics. Suno or Udio cover the vocal-song territory; in practice many producers pair the two approaches, pulling the instrumental from Stable Audio and adding a voice from another tool or a real microphone.
What hardware does it take to run locally?
Very little for Small: it works without a GPU at all, through CoreML on Mac, and composes offline on a phone. Medium wants a graphics card with roughly 5 to 6.5 GB of video memory depending on the decoding mode. Large is not distributed as open weights and stays server-side only.
Verdict: No sung chorus, and that is deliberate: video editors, game studios and producers hunting for instrumentals, ambient beds and sound effects with a traceable training history get, in Stable Audio 3.0, an engine to own rather than rent.
