Contents

42 / 153

Video Models, by Company

Google DeepMind (Veo)

Chapter 41

11 min read

Reviewed v78 · August 2026

01

Veo's strategic position

Veo is Google DeepMind's video generation model. It was announced in May 2024 at Google I/O, three months after Sora's preview, as Google's response to OpenAI. Where Sora had to be an independent product because OpenAI's strategic position required it, Veo was designed from the start to be integrated into Google's existing creative tools, YouTube (eventually), Vids (Google's video creation app), Slides, the Gemini app, and the Google Workspace productivity suite. Veo is what shows up when you click 'generate video' inside Google's products, and that distribution gives it a massive built-in user base that no startup can match.

Veo, like Imagen and Nano Banana, benefits from Google's effectively unlimited compute and unique data resources. YouTube alone is the world's largest collection of video content, and while we do not know exactly what subset Google uses for training, the existence of that resource shapes what is possible. The model is clearly trained on enormous amounts of video, and the results show it, Veo produces some of the most cinematically convincing outputs in the field, with strong understanding of camera language, lighting, and motion.

02

Veo 1, Veo 2, the early generations

Veo 1 was announced in May 2024 and integrated into Google's products through 2024. Quality was competitive with Sora's preview and somewhat better than the publicly released Sora when it eventually came out in December. Veo 2 launched in late 2024 with significant quality improvements, better prompt following, more physically realistic motion, longer clip durations. These generations were good enough to be useful but not yet at the level that made the field universally impressed.

03

Architecture and core ideas

Veo 3 and 3.1 are described in Google's technical communications as latent diffusion transformers built around a 3D variational autoencoder that compresses video and audio into a joint latent representation. Three architectural choices distinguish Veo from the other frontier video models. First, the VAE compresses both visual and audio content into a single shared latent space rather than treating them as separate modalities, which means the diffusion process is genuinely joint rather than sequential. When Veo generates a scene of a door slamming, the visual frames and the sound of the slam are sampled from the same denoising trajectory, not stitched together after the fact. This is why Veo's audio synchronization is so tight, the model is not learning to match sound to image, it is learning to generate them as one object.

Second, Veo uses the 'spatiotemporal patches' approach that Sora pioneered, treating video as a three-dimensional tensor (width, height, time) that gets tokenized into small patches and fed to a transformer. This is the architectural pattern most of the 2024-2025 video model generation converged on, and Veo's implementation is the most polished version of it shipping in production. Third, Veo's text encoder draws directly on Google's internal language model stack (Gemini), which gives the model much stronger world knowledge and cinematographic vocabulary than a model trained only on video-caption pairs. When you write 'dolly in, rack focus, low angle, volumetric lighting,' Veo understands what each of those terms means because the same model backbone has absorbed millions of film textbooks, YouTube tutorials, and cinematography discussion threads from Google Search's training data.

04

Veo 3 and Veo 3.1 (May 2025 and October 2025)

Veo 3, announced in May 2025 at Google I/O, was the breakthrough release. Two things made it stand out. First, native joint audio-video generation. Veo 3 was, alongside Sora 2, one of the first widely accessible models that could produce video with synchronized audio in a single generation step, including dialogue, sound effects, and ambient noise. Second, the video quality itself was a clear step forward, Veo 3 outputs at 1080p (and Veo 3 with upscaling can produce 4K content) with motion that consistently looks plausible rather than uncanny.

Veo 3.1, released in October 2025, was the refinement release that made Veo the current SOTA leader in video generation. Native 48 kHz audio replaced the lower-bitrate audio in Veo 3, bringing sound quality to broadcast standard. Native 4K output became standard rather than an upscaling post-process. Prompt understanding was significantly improved through better text encoder integration, and new creative controls were added including camera path specification and explicit motion intensity parameters. Pricing through the API is in the range of $0.40 to $1.50 per second of generated video depending on tier and resolution, which makes Veo one of the more expensive per-second video models but also the one that most production teams route to for hero shots where quality matters more than cost.

As of July 2026, the standings have shifted. On the public arenas Google's Gemini Omni Flash now leads overall video quality, with Veo 3.1 still a top-tier cinematic option, Kling 3.0 as the strongest multi-shot storyteller, ByteDance's Seedance line as the character-consistency and long-take leader, and Wan as the strongest open-weight option. Most production teams route between these based on scene type. Veo 3.1, and increasingly Gemini Omni Flash, is the default for cinematic hero shots, for any scene where audio quality has to be broadcast-ready, and for any work that needs to hit 4K deliverables without a separate upscaling step.

05

Strengths and weaknesses

Veo's strengths are cinematic quality, audio-video joint generation, native 4K, and the underlying world knowledge that comes from being built on top of Gemini's text understanding stack. For finished cinematic work where visual polish matters more than anything else, Veo 3.1 remains a top choice for finished cinematic work as of mid-2026, even as Google's Gemini Omni Flash has taken the top of the public video arenas. The audio integration is particularly unmatched, other models generate plausible audio but Veo's audio is genuinely broadcast-quality, with proper dialogue lip sync, accurate Foley for actions visible in the frame, and ambient sound that matches the environment. Google's cinematographic training data also gives Veo the deepest understanding of camera and lighting terminology of any video model, prompts written in the vocabulary of professional film production produce results that match what a working cinematographer would expect.

Veo's weaknesses are cost, access constraints, and the same closed-model customization problems that affect the rest of Google's image and video lineup. Per-second pricing is among the highest in the market, which makes Veo expensive at any meaningful volume. Access is through Google's products (Vertex AI, the Gemini API, Google Vids, Google Workspace integrations) or through a small number of partner platforms (Krea, fal.ai, Replicate), and the API terms are less flexible than the open-weight alternatives. There is no fine-tuning, no LoRA support, no offline use, and no weights available. For custom brand work where you need to condition generations on proprietary character designs or product libraries, Veo currently has no good answer, you have to fall back on the last open-weight Wan release, Wan 2.2, for that kind of work. Veo's motion coherence also still degrades above roughly 15 seconds of clip length, which is better than most competitors but not yet at the multi-minute coherent storytelling that would unlock true long-form narrative video.

06

Strategic position

Veo's strategic position is the mirror image of Midjourney's. Where Midjourney has a devoted consumer creative base and no infrastructure reach, Veo has massive distribution through Google's products but no consumer creative brand. Most users who encounter Veo experience it as 'the video generation feature inside Google Vids' or 'the thing that makes videos when I click the magic button in Google Slides,' and most of them do not know or care what model is producing the output. This has been a deliberate choice by Google, video generation is a feature of the Google productivity stack rather than a standalone product, which is the same pattern that DeepMind follows for Gemini image generation through Nano Banana.

For operators, Veo is the clearest case of a top-tier frontier model that you have to decide whether to use based on cost rather than capability. The capability is there, the quality is there, the audio integration is unmatched. The question is whether your unit economics support per-second pricing that can reach $1.50 for hero shots, and whether the walled-garden access model works for your pipeline. Teams building high-volume video products typically route only their highest-value generations to Veo and use cheaper alternatives (Wan 2.6, Hailuo, or a distilled variant) for the bulk of production. Teams building premium creative tools often use Veo as the default because the per-generation quality justifies the cost. The Sora shutdown covered in the OpenAI Sora chapter also left Google in a structurally stronger position, since Sora was Veo's most direct strategic competitor and the retreat of that competitor gives Google more room to set pricing and pace of release on its own terms.

07

Gemini Omni Flash and the end of the standalone Veo line

On June 30, 2026, Google announced Gemini Omni Flash, a fast and cost-efficient model for both video generation and conversational, chat-based video editing. It takes text, image, audio, and video as input and produces video with synchronized audio, and it is available through Google AI Studio, the Gemini API, and Google Flow. The 'Flash' naming is the same signal it carries elsewhere in Google's lineup: this is the cheap, quick tier meant to run at volume, not a slow prestige model.

That volume tier is now, by the public numbers, the best video model in the world. As of July 2026 Gemini Omni Flash ranks first on both the Artificial Analysis text-to-video and image-to-video arenas, ahead of ByteDance's Seedance line, Alibaba's Wan, and the Kling 3.0 variants. It is worth pausing on that: the model that took the top of the leaderboard is Google's fast, low-cost branch, not a separately marketed flagship. The Veo product page still lists Veo 3.1 as the current standalone version, with eight-second clips, 1080p and 4K options, native audio, character consistency, scene extension, and first and last frame control. But the strategic center of gravity has clearly moved to the unified Gemini Omni assistant, where video is one capability among many rather than its own product.

08

The org behind Veo, and the moat

Veo is built by the generative-media team inside Google DeepMind, the merged research organization run by Demis Hassabis, the same group that produces the Imagen image models and the Lyria music models. The reason to care about the org chart is that it explains the strategy. Almost no competitor controls the whole stack the way Google does. It owns the research lab, it owns the training compute in the form of in-house TPU pods rather than rented Nvidia clusters, it owns a video corpus of staggering scale through YouTube, and it owns more distribution surfaces than anyone: the consumer Gemini app, the Flow filmmaking tool, Vertex AI for enterprise, the Gemini API for developers, and YouTube Shorts. When one company holds the lab, the chips, the data, and the shelf space at once, its costs and its reach are simply different from a startup renting all four.

Hassabis is also the clearest voice on where this is going, and he has been consistent. He called Veo 3's native, synchronized audio the moment AI video left the era of the silent film, and back in April 2025 he said plainly that Google would eventually fold Veo into Gemini rather than keep it a separate product line, a prediction that arrived as Gemini Omni. The other tell about ambition is the money moving outward: in mid-2026 Google put a reported seventy-five million dollars into the independent studio A24 to co-develop filmmaking tools, and notably structured the deal so Google does not get access to A24's content library. That is a partnership aimed at creative legitimacy and workflow, not a data grab, which is a different and more patient bet than most of the field is making.

09

Getting the best out of Veo

There are four doors into the same models, and picking the right one matters. For hands-on creative work, Flow at labs.google is the dedicated filmmaking studio, with a scene builder, camera controls, and reference images it calls ingredients for holding a character or a look steady across shots. For casual generation the Gemini app bundles video into the assistant. For production with governance and service levels, Vertex AI on Google Cloud is the path. For programmatic use, the Gemini API and AI Studio expose everything. Flow meters generation with credits across the Google AI subscription tiers, running from a limited free allowance through a Pro plan around twenty dollars a month up into an Ultra band that costs substantially more for the largest credit pools.

On the API the pricing is per second and tier-dependent: the top Veo 3.1 tier runs on the order of forty cents a second at standard resolution and around sixty cents at 4K, while the Fast and Lite tiers drop that by roughly three to eight times for drafts and volume work, and you are billed only on a successful generation. As of mid-2026 the practical default is Gemini Omni Flash for general text, image, or audio driven video and for conversational, multi-turn editing, and you reach for Veo 3.1 specifically when you need scene extension, precise first or last frame control, or an existing pipeline built on a Veo model ID. On prompting, be a cinematographer in words: name the camera move, the lens, the lighting, and the motion explicitly, lean on reference ingredients for consistency, and because the audio is generated natively, actually write out the dialogue, the sound effects, and the ambience you want to hear. It is strongest on short cinematic shots with matched sound and believable physics, and weakest on long single takes, dense on-screen text, and anything the safety filters flag.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. How does Veo's strategic positioning differ from Sora's?

  2. 2. What made Veo 3 a breakthrough release?

  3. 3. Why is Veo's audio synchronization so tight?

  4. 4. Why does Veo understand cinematographic terms like 'rack focus' and 'volumetric lighting'?

  5. 5. What did Veo 3.1 add over Veo 3?

  6. 6. What is a stated weakness of Veo for custom brand work?

  7. 7. Why does Google ship a distilled version of Veo rather than its full internal model?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.