Suno and Udio get the attention, but the music and sound field is broader, and the interesting split is not who sounds best this week, it is who trained on what. On one side sit players built on licensed or owned data, selling commercial safety. On the other sit open models you can run yourself. And looming over all of it is the sound now being generated inside the video models themselves.
The rest of the music field
Google's Lyria line is less a single product than the music engine threaded through Google's whole ecosystem. Lyria 2, opened up in 2025, powers the MusicFX generator, the professional Music AI Sandbox, and real-time tools, and it feeds YouTube's Dream Track for Shorts. By 2026 a newer Lyria generation reached the Gemini app, turning text or images into short songs. Every output is watermarked with SynthID, the same provenance layer Google says has now marked over 10 billion pieces of media across image, video, audio, and text. Google's real edge here is not raw musical quality but distribution and reach through YouTube, though its consumer clips have tended to stay short.
The clearest sign of how the field is consolidating is what became of Riffusion. The project that began as a viral experiment generating music from spectrogram images had rebranded to ProducerAI, a browser-based tool that assembles full songs and even custom instruments, and in February 2026 Google acquired it and folded the team into Labs and DeepMind, with ProducerAI running on a preview of the newer Lyria. In roughly three years, a scrappy open experiment became an acquisition target for the company building the platform's music engine.
Stability's Stable Audio is the purest licensed-data play. Stable Audio 2.5, released in September 2025, was pitched explicitly at enterprise sound production, trained on a fully licensed dataset with legal indemnification for business customers and fast enough to render a multi-part track in seconds. Stable Audio 3.0 followed in 2026, generating tracks up to about six minutes and, notably, shipping most of its model tiers as open weights, with the largest reserved for API and enterprise use. Stability trained it entirely on licensed data, supported by partnerships with major labels, and made commercial use free below a revenue threshold. The whole proposition is that you can ship the output without a lawyer's call.
The open research foundation much of the community still builds on is Meta's AudioCraft, which bundles MusicGen, AudioGen, and the EnCodec codec. Its licensing carries a subtle but important catch: the code is MIT-licensed, but the model weights are released under a non-commercial license, which makes AudioCraft a research and prototyping tool rather than something you can ship a commercial product on. The open frontier has since moved partly to China. Tencent AI Lab open-sourced its SongGeneration model, and its 2026 successor is a roughly four-billion-parameter model that generates complete songs several minutes long, alongside other open efforts such as ACE-Step. Music generation is not a Western-only race, and some of the strongest fully open models now come from Chinese labs.
Sound effects and foley
Sound effects are the quiet workhorse of this field. ElevenLabs, Stability, and others generate effects from a text prompt and return several candidates to choose from, and the more useful variant for film is video-to-sound: point the model at a clip and have it generate foley timed to the on-screen action, so an impact lands on the exact frame of contact. ElevenLabs shipped a video-to-sound feature that analyzes a clip and lays in matching effects, and extended its sound-effects model toward longer clips with seamless looping. On the open side, research systems like MMAudio tackle the same video-to-audio synchronization problem. This is exactly the problem the video models are now solving natively, which is why dedicated foley tools are converging with, and competing directly against, native audio in video.