Contents

150 / 153

Additional Notable Models and Tools

Audio and music

Chapter 149

1 min read

Reviewed v78 · August 2026

Audio generation is its own large field that is somewhat outside the scope of this document, but a few models are relevant because they integrate with image and video workflows.

01

ElevenLabs

The dominant text-to-speech and voice cloning service. Used for adding dialogue and narration to AI-generated video. Their voice cloning is high enough quality that most casual viewers cannot tell it apart from real recordings.

02

Suno and Udio

Two competing music generation services that produce full songs from text prompts, including vocals. Frequently used for adding soundtracks to AI-generated video.

03

Stable Audio

Stability AI's audio generation model. Generates sound effects and short audio clips from text. Useful for matching audio to specific moments in generated video.

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.