If one company defines AI audio the way OpenAI defined chat, it is ElevenLabs. It is the quality leader in voice, the most widely used, and, having expanded from text-to-speech into cloning, dubbing, sound effects, music, transcription, and voice agents, it is trying to become the audio layer for the entire AI era rather than a single tool.
The company: two friends and a decade of bad dubbing
ElevenLabs was founded in 2022 by two childhood friends from Poland: Piotr Dabkowski, an ex-Google machine-learning engineer who is CTO, and Mati Staniszewski, an ex-Palantir strategist who is CEO. Their origin story is unusually on-point for the product: they grew up watching American films clumsily dubbed into Polish, where the dubbing flattened the emotion of the original performances, and they wanted to fix it. Rather than ship a quick chatbot wrapper at the peak of the 2022 hype, they spent roughly a year in stealth getting the voice quality right, which turned out to be the correct instinct.
The funding ladder tells the rest of the story. A 19-million-dollar Series A in 2023, an 80-million-dollar Series B in January 2024 at over a billion, a 180-million-dollar Series C in January 2025 at 3.3 billion, and a 500-million-dollar Series D in February 2026 at an 11-billion-dollar valuation, a roughly hundredfold increase in under three years. The company reported annual recurring revenue over 330 million dollars at the end of 2025 with only around 350 employees, an unusually lean, high-revenue-per-head operation. It organizes itself into three lines that are worth keeping straight: the conversational-agent business, the creator business, and the developer API. The agent business is probably the biggest growth engine, but the creator and API lines are what matter for filmmaking.
The models and the product surface
The model line is a spectrum from quality to speed. Multilingual v2, from 2023, covers nearly thirty languages and is still positioned as the most lifelike option for long-form work like audiobooks, at the cost of latency. The Turbo and Flash models trade some expressiveness for speed, with Flash generating speech in about 75 milliseconds, which is what makes real-time voice agents feel live. The expressive flagship is Eleven v3, and its signature idea is worth calling out: you direct the performance with inline tags in the script, writing things like a whisper, a laugh, or a sarcastic delivery in brackets, and a Text to Dialogue mode that manages turns and interruptions across multiple speakers. It is effectively a scripting language for emotion.
Around the models is a full audio stack, which is the real reason ElevenLabs matters to a filmmaker. Instant Voice Cloning builds a usable clone from about a minute of audio; Professional Voice Cloning trains a higher-fidelity one from thirty-plus minutes and requires the speaker to verify their identity. Voice Design generates a brand-new synthetic voice from a description, which sidesteps likeness issues entirely. Dubbing transcribes, translates, and re-voices content across languages while trying to preserve the original timbre. Studio is the long-form workspace for assembling dialogue, music, and effects on a timeline. And the company has pushed beyond voice into a Sound Effects model, a music generator called Eleven Music trained under opt-in label licenses, and a Scribe speech-to-text line. One honest caveat for filmmakers: the standalone Dubbing Studio is described in the docs as being in maintenance mode, so serious dubbing pipelines lean on the dubbing API plus manual cleanup rather than expecting new features there.
Getting the best, and holding a voice across a film
Access is credit-based, from a free non-commercial tier up through creator and pro plans that unlock professional cloning, commercial rights, and higher audio quality, with the same credit system exposed through the API at a lower per-unit cost on the fast Flash model. Verify the exact numbers on the live pricing page, because they move. For the best output, reach for Multilingual v2 or v3 rather than Flash when quality matters, give the model clean well-punctuated text, and use the v3 emotion tags sparingly and test them per voice, because they do not fire reliably on every voice.
The hardest practical problem in a film is keeping a character's voice consistent across pickups, reshoots, and scenes shot months apart, and it does not happen by accident. The reliable pattern is to build one Professional Voice Clone or one designed voice per character, save and reuse the exact voice settings, keep the same model version throughout, since mixing versions shifts the timbre, and regenerate line by line so a single bad take does not force re-rendering a whole scene. For serious narrative work the field increasingly does not use raw text-to-speech at all: it uses speech-to-speech conversion, where a real actor performs the emotion and timing and the tool only swaps the voice identity, which we cover in the voice-landscape chapter. ElevenLabs is strongest at expressive English narration, character dialogue, and real-time agents, and weakest at guaranteed emotion-tag reliability and at cost efficiency for very high-volume dubbing.
The likeness problem
Voice cloning is both the core capability and the core liability, and ElevenLabs sits at the center of the field's hardest ethical question. In January 2024 its tools were forensically tied, with over 99 percent confidence, to a fake robocall of President Biden telling New Hampshire voters not to vote, which became the defining example of AI voice used for election disinformation. The company suspended the account, and it has built real safeguards, a voice-cloning captcha, identity verification for professional clones, a speech classifier, and bans, but the safeguards are as much reactive as preventive.