A guide to a medium that now speaks cannot stop at pixels. Audio is nearly half of the production phase. Voice and dubbing is led by ElevenLabs, with a deep field behind it (Play HT, Cartesia, Hume, Respeecher) that generates, clones, and localizes speech, while lip-sync and avatar tools like HeyGen, Synthesia, Hedra, and Sync marry a voice to a face. Music generation, led by Suno and Udio with ElevenLabs and Stable Audio close behind, produces full songs from a prompt, and a smaller sound-effects layer, ElevenLabs again and the open MMAudio, fills in foley and ambience. As the foundations noted, the frontier video models now generate synchronized audio natively, so the line between the video and audio pillars is starting to blur at the top.
Beyond sound, three newer categories are worth holding in view. Three-dimensional generation, from Meshy, Tripo, Luma's Genie, Rodin, and Tencent's open Hunyuan 3D, turns text and images into assets and scenes for games and AR. World models are the most ambitious box on the board, and in 2026 the fastest-moving: rather than flat clips they generate explorable, persistent environments you can move through. Google DeepMind's Genie 3, Fei-Fei Li's World Labs and its Marble product, NVIDIA's Cosmos, Decart, Odyssey, and Runway's own effort are all chasing it, and the money has followed. Reactor came out of stealth in May 2026 with a 59 million dollar round, led by former Apple Vision Pro engineers, to build a developer platform for real-time AI worlds, and Tencent shipped an open real-time world model, HY-World, that streams at interactive frame rates. And a programmatic layer, from Remotion's code-rendered video to agentic chat interfaces, is quietly turning generation into something software calls on its own. Two adjacent categories, AI ads (Arcads, Icon) and interactive characters (character.ai, Fable's Showrunner), point at where the money and the audiences already are.
Speech, cloned and carried across languages
This category turns text into speech and moves that speech between voices and languages. It exists because generated video is silent by default, and because the hardest part of localization was never the translation, it was making a performance sound like the same person speaking a new tongue. The work splits into three overlapping jobs: synthesizing natural speech from a script, cloning a specific voice so it can say anything, and dubbing existing footage so a film travels without losing its actors.
ElevenLabs is the category's center of gravity, the leading voice synthesis and dubbing platform that has since pushed outward into music and sound effects. Around it sit labs betting on latency and architecture: Cartesia builds low-latency state-space voice models under the Sonic name, an approach aimed at voice you can talk to in real time, while Play HT focuses on realtime synthesis and voice agents for the same conversational use. Hume takes a different angle with its EVI interface, treating emotional expressiveness as the product rather than a side effect.
The cloning and dubbing side is where fidelity matters most. Respeecher offers high-fidelity voice cloning trusted in film and games, the kind of work where a cloned voice has to survive close listening. Deepdub and Dubverse both handle dubbing and localization for video, turning one language track into many. On the research edge, Kyutai is an open speech lab whose Moshi model pushed full-duplex conversational voice into the open, and Fish Audio leans open with its TTS and cloning. Google's NotebookLM sits here too, less as a voice engine than as the tool that popularized the AI audio overview, a two-host conversation generated from your own documents.
The talking head, manufactured
Avatar tools generate a person on camera who never stood in front of one. The category exists because a huge share of practical video is simply someone talking: explainers, training, marketing, localized announcements. If you can drive a face from a script and a voice, you can produce that footage without a shoot, and you can reshoot it in forty languages by swapping the audio. The two technical problems are building a believable presenter and getting the mouth to match the words, which is why lip-sync and avatars live in the same category.
HeyGen and Synthesia are the two names most teams reach for. Synthesia built the enterprise standard for avatar video with multilingual narration, aimed squarely at corporate training and internal comms, while HeyGen covers similar talking-head and marketing ground with a faster, more consumer-facing feel. Captions comes at the same market from mobile first, packaging avatars, captions, and AI video editing into a phone-native app.
Underneath the polished presenters is the raw capability, and Sync.so exposes it directly as a lip-sync API that matches any voice to any face, useful when you already have footage and only need the mouth to agree with new audio. Hedra pushes toward expressive character video rather than neutral presenters, and D-ID has long specialized in the talking-photo, animating a single still image into a speaking face. LemonSlice is chasing the realtime version of the same trick, a talking avatar that responds live rather than rendering after the fact.
A song from a sentence
Music generation produces finished audio, often a full song with vocals, from a text prompt. It exists because scoring and licensing music has always been slow and expensive for anyone making video at volume, and because a model that understands both composition and production can hand a creator something usable in a minute. The category spans two poles: consumer song generators built for delight, and quieter tools built to score a video without ever putting a hook in your head.
Suno and Udio are the frontier, the two systems that made prompt-to-song feel real. Suno is the more consumer-famous of the pair, tuned for full text-to-song generation, while Udio competes on fidelity. Both operate under the music industry's close and often litigious attention, which shapes how they talk about training data and output. Stable Audio, from Stability, takes the open-leaning path and emphasizes long-form generation, and ElevenLabs Music extends the voice leader's reach into songs.
The rest of the category is built for utility rather than hits. Aiva is composition-focused and aimed at scoring, closer to a film composer's assistant than a jukebox. Beatoven.ai, Mubert, and Soundraw all sell the same practical promise in slightly different shapes: royalty-free, often adaptive music that a video editor can drop in without a licensing headache, with Mubert leaning on generative streams and an API and Soundraw on customizable royalty-free tracks.
Foley without the prop room
Sound effects generation makes the footsteps, door slams, wind, and impacts that sell a shot. It is a small category with an outsized job, because generated video arrives silent and the human ear notices missing sound long before it notices a flawed frame. The interesting technical split is between prompting a sound from text and inferring it from the video itself, which is the difference between describing what you want and letting a model watch the picture and score it.
ElevenLabs SFX is the straightforward text-to-sound-effects entry, part of the same house expanding out from voice. The more research-forward idea is video-to-audio foley, and MMAudio represents that open line of work, generating sound conditioned on the moving image so the audio lands on the action rather than near it. Mirelo works the sound-effect generation problem as well.
The category also reaches toward music-adjacent production. ACE Studio focuses on AI vocals and sound production, and Adobe Firefly Audio brings generative audio into Adobe's commercially cautious ecosystem, which for many teams is the deciding factor when the output has to clear a brand's legal review.
Meshes on demand
3D production tools generate assets and scenes from text or images, meshes and textures rather than flat pictures. The category exists because 3D content is expensive to model by hand and is needed in enormous quantity by games, AR, virtual production, and product visualization. A tool that turns a prompt or a single photo into a usable mesh collapses hours of modeling into a first draft, even if a human still cleans it up.
Meshy is a popular general-purpose generator for text and image to 3D, widely used by game and AR teams, and Tripo competes on fast image-to-3D mesh output. Rodin, from Hyper3D, aims at higher-quality asset generation, and Tencent's Hunyuan 3D brings an open option into the mix, part of the broader pattern of Chinese labs shipping open weights. Luma's Genie extends that studio's generative work into text-to-3D.
Around the pure generators sit tools shaped by specific pipelines. Kaedim builds 2D-to-3D production models aimed at game studios, Sloyd does parametric real-time asset generation for cases where speed and editability beat raw fidelity, and Spline offers web-native 3D design with AI woven in. At the far end, Simulon pushes photoreal VFX and virtual production onto a phone, and Autodesk Maya represents the incumbent path, the industry-standard package adding AI assists rather than being replaced by them.
Worlds you can walk into
World models generate environments you can move through, not clips you watch. The distinction that defines the category is interactivity and persistence: the world responds to your input and, ideally, remembers what it looked like when you turn back around. This is the frontier where generated video, game engines, and robotics simulation converge, and it is early, expensive, and moving fast.
Google DeepMind's Genie 3 is the headline, generating realtime, interactive, persistent 3D worlds, and it sets the bar the rest of the field measures against. World Labs, Fei-Fei Li's spatial-intelligence lab, is the most watched independent effort and has moved from research into product with its Marble launch. NVIDIA Cosmos comes at world models from the robotics and simulation side, offering a foundation for training machines rather than entertaining people, which is a reminder that this category serves two very different customers.
The realtime, playable end is crowded with fast movers. Decart builds realtime playable world models under the Oasis name, Odyssey is a well-funded entrant pursuing interactive world and video models, and Meta's WorldGen represents the largest platform's own research bet. Runway, better known for video, is also pursuing general world models, a sign of how naturally the video labs drift toward this problem once their clips get long and controllable enough.
Generation by conversation or by code
This category is about how you drive generation rather than what generates. It splits cleanly in two: agentic chat, where you talk to an interface that creates and edits media for you, and programmatic, where video is produced by code and data with no human in the loop at all. Both exist to remove the manual timeline. One replaces it with conversation, the other with an API call.
On the conversational side, Luma's chat interface lets you generate and edit media by talking to it, and OiiOii pursues the same agentic-chat approach to media creation. The appeal is that the model handles the tool operation while you stay in plain language.
The programmatic side is where automation at scale lives. Remotion is the standout idea, rendering video from React code, which lets engineers build video the way they build web pages. Creatomate and JSON2Video take the templated route, generating data-driven video from a spec so a single design can be rendered a thousand times with different content. Shotstack offers a cloud video-editing API built for the same automation, and Cloudglue works the video-understanding and generation side through an API. Together these tools power the personalized and bulk video that no team could hand-edit.
Creative built for the feed
This category applies generative media to the specific job of performance marketing and social content, where the goal is not one perfect film but hundreds of variants tested against an algorithm. It exists because ad creative is consumed and discarded at a rate no traditional production process can feed, and because the format rewards volume and iteration over polish. These tools are built to produce, test, and refresh creative continuously.
Arcads is a clear expression of the idea, generating UGC-style ads with avatars at scale, manufacturing the casual talking-to-camera format that performs on social. Icon, under the Mirage name, automates performance-ad generation directly, and Bluma and Zeely work the same ad-creative ground, with Zeely aimed at the SMB end where a small store needs ads and storefront content without an agency.
The enterprise and copy-first side is anchored by two established names. Typeface focuses on brand-safe content generation for larger companies, where staying on-brand and out of legal trouble is the whole point, and Jasper is the marketing copy and content platform that many teams already run. At the more experimental edge, eggnog builds AI creators and character-driven social video, betting that recurring synthetic personalities, not just single ads, are the unit of attention.
Stories that talk back
The Interactive category covers experiences where the audience participates rather than watches, playable characters and generative narratives that respond to input. It exists because generative models make it cheap to produce content on demand, which turns a story from a fixed artifact into something that can be regenerated around each user's choices. This is the point where entertainment stops being a file and becomes a session.
character.ai is the largest example, a platform of interactive AI characters for chat and roleplay that reaches a very large and often young audience. Cantina and Shizuku AI pursue related visions of AI character social platforms and experiences, treating synthetic personalities as something you hang out with rather than prompt once.
The narrative-generation side aims higher up the story stack. Showrunner, from Fable, generates entire AI TV episodes, an ambitious attempt at a fully synthetic show. Lore Machine turns a story into an illustrated narrative, Wide Worlds builds interactive AI narrative worlds, and Pickford works the interactive storytelling ground. The common thread is a story engine that the reader can steer, which is a very different product than a finished film even when it borrows the same underlying models.