Contents

43 / 153

Video Models, by Company

Kuaishou (Kling)

Chapter 42

10 min read

Reviewed v78 · August 2026

01

The Chinese answer to Sora

Kling is the most successful and most consistently improving video model from China. It is built by Kuaishou, a publicly traded company on the Hong Kong Stock Exchange that operates the second-largest short-video platform in China (after Douyin, the Chinese version of TikTok). Kuaishou has hundreds of millions of users and a large internal AI research team, and Kling is the flagship product of their generative AI work.

Kling was unveiled in June 2024, four months after Sora's preview, and at a moment when most observers assumed OpenAI's model was unrivalable. The first version of Kling immediately demonstrated quality in the same league as Sora's preview, with the additional advantage that it was actually accessible to users (initially through Kuaishou's Chinese-language app, and later through a Western web interface). The release was a wake-up call for the field: the gap between American and Chinese AI labs in video generation was much smaller than people had assumed, and in some respects Kling was already ahead.

02

Architecture

Kling uses a Diffusion Transformer architecture combined with a custom 3D Variational Autoencoder that Kuaishou built in-house. The 3D VAE performs what they call 'synchronous spatiotemporal compression', encoding both spatial information within frames and temporal information across frames simultaneously. This is a meaningful technical claim: rather than encoding frames independently and then learning temporal relationships between them, the VAE bakes time into the latent representation from the beginning. The result, Kuaishou says, is improved temporal coherence at the most fundamental level of the pipeline. Treat those specifics with some care, though, because the core Kling video model has never shipped a technical paper, so the 3D VAE and its compression claims come from the company's own materials rather than published work. What is verifiable comes from the framework papers Kuaishou has released around the model: the Kling-Omni report describes the current stack as a diffusion transformer aligned with a vision-language model, and Kling-Foley is a six-billion-parameter multimodal diffusion transformer that generates sound aligned frame by frame to the video. The honest summary is that the shape is confirmed and the exact internals are not.

Kling also uses a 'computationally efficient full-attention mechanism' for the temporal modeling, meaning the transformer attends across all spatial and temporal patches together rather than splitting attention into separate spatial and temporal phases. This is more expensive than factored attention but produces better motion coherence, and Kuaishou has clearly decided the quality is worth the cost.

03

Kling versions

Kling 1.0 (June 2024)

The original public release. Generated up to 2-minute clips at 1080p and 30fps. Already competitive with Sora's preview at launch. Generated significant attention in the AI community as proof that the gap to OpenAI was small.

Kling 1.6 (December 2024)

Refinement with improved video quality. Established Kling as a serious commercial competitor and started to attract Western users in addition to Chinese ones.

Kling 2.0 (April 2025)

Major upgrade. Better motion, better prompt following, better physical realism. The 2.x line is where Kling really matured into a top-tier model.

Kling 2.1 (May 2025)

Introduced selectable quality modes, letting users trade off generation time against output quality. Also brought significant improvements to camera control, Kling 2.1 understands cinematographic terminology and can execute requests like 'dolly zoom,' 'tracking shot,' and 'crane down' with the precision a director would expect.

Kling 2.5 (mid-2025)

Continued refinement and feature additions, including better handling of human characters and improved consistency across multi-shot sequences.

Kling 3.0 (February 7, 2026)

The current flagship as of mid-2026. Kling 3.0 is described as a unified multimodal model that generates video, audio, and images within a single architecture, meaning that, like Sora 2 and Veo 3, it now does joint audio-video generation natively. The feature set has expanded to include 15-second cinematic generation, native multi-shot storyboarding (the model can generate sequences of related shots that maintain character and scene consistency), and significantly improved physical realism. Kling 3.0 is a strong top-tier model, though the honest read as of mid-2026 is that the public arenas place it just behind the leaders, Google's Gemini Omni Flash, ByteDance's Seedance, and Alibaba's Wan, rather than tied for the very top.

Kling's annualized revenue run rate climbed from roughly $100 million in early 2025 to about $500 million by March 2026, and by mid-2026 Kuaishou reported more than 100 million users across 224 countries and regions, over 600 million videos generated, and around 50,000 enterprise customers. It is one of the most commercially successful generative video products in existence.

04

The spin-off and Kling 3.0 Omni

The product kept moving alongside the corporate maneuvering. The base Kling 3.0 model shipped in February 2026, and a Kling 3.0 Omni upgrade and a cheaper Turbo variant with bundled audio followed in the middle of the year, adding stronger cross-shot consistency and 4K editing. The exact dates on the mid-year Omni release come from secondary trackers rather than a Kuaishou press release, so treat the specifics as provisional, but a Kling 3.0 Omni entry does appear on the public arenas, which corroborates that the model exists and is competitive.

05

The business behind the model

It is worth understanding who is actually building Kling, because it explains a lot about the product. Kuaishou is China's second-largest short-video and live-streaming company, the main domestic rival to ByteDance's Douyin, founded in 2011 by Su Hua and Cheng Yixiao and listed in Hong Kong in 2021. Cheng Yixiao is now CEO, and the Kling AI unit is run by Gai Kun, a senior vice president who reports to him; in mid-2026 Kuaishou moved more senior technical leaders into the unit ahead of the spin-off. A short-video company builds a frontier video model for the same reason ByteDance does: it already owns an enormous proprietary corpus of video and a built-in audience to distribute and monetize the output.

That advantage is now being priced. In July 2026 Kuaishou carved Kling into its own subsidiary and raised roughly 2.8 billion dollars, a valuation near 18 billion, keeping about 68 percent, with Tencent, Alibaba, and Baidu among more than thirty investors, and a clause that triggers a buyback if Kling has not gone public by 2031. The standalone numbers are those of a fast-growing, still-unprofitable business: on the order of 1.1 billion yuan of revenue in 2025 against a wider loss, first-quarter 2026 revenue up more than threefold year on year, and an annualized run rate near 500 million dollars. The reach is real: over 100 million users across 224 countries, hundreds of millions of videos generated, tens of thousands of enterprise customers.

06

What Kling is best and worst at

Kling's signature strength has always been motion. Its clips have a cinematic sense of weight and camera movement that many rivals still struggle to match, and that shows up in adoption: on the public API platforms, Kling's image-to-video models are among the most-run video models anywhere, which is a real signal of working-creator trust rather than benchmark noise. Version 3.0 added the things it most lacked, native audio with lip-synced dialogue, multi-shot sequences of up to six shots with per-shot prompts, and up to seven reference images to hold a character and style across a generation, plus a dedicated motion-transfer model that maps the movement of a reference video onto a still character, a first-class feature many competitors do not have.

The honest weaknesses are just as clear. On the mid-2026 arenas Kling sits around sixth, genuinely strong but behind Google's Gemini Omni Flash, ByteDance's Seedance, and Alibaba's Wan, and its flagship is priced above Seedance and Wan while ranking below them, which is an uncomfortable place to be. The new native audio is caveated: it works best in English and Chinese, and it cannot be used together with a reference-video input, which breaks some workflows. Clips still cap around fifteen seconds. Complex physics interactions, in Kling's own words, may not look fully natural, and a character can drift in appearance across separate generations even when it holds within one. And because it is a Chinese-platform product, its content filters are strict, politically sensitive material is blocked, which matters if that is your subject.

07

Getting the best out of it

Access comes in a few shapes. There is an international web app at kling.ai and a China-domestic version at klingai.com, the Kling mobile app, exposure through Kuaishou's Kwai app, an official developer API sold as prepaid credit packs, and the full lineup on the third-party inference platforms, fal and Replicate among them, which is usually the easiest route for a developer. Within the product, the quality modes are the first lever: Standard runs at 720p and is cheaper and faster, Pro runs at 1080p and unlocks the end-frame control, and the Master tier from the 2.x line is the highest-fidelity, highest-cost option. Subscriptions run from a free daily-credit tier up through Standard, Pro, Premier, and Ultra, and on the API the current models price roughly between eight and twenty cents per second depending on tier and whether audio is on.

The prompting rewards a specific discipline. Describe subject, then motion, then one camera move, then lighting and style, and describe the motion physically, arms swinging, heel-first steps, weight transfer, rather than vaguely. Use one camera movement per clip, because stacking several confuses the model. Always give the motion an endpoint, hair moves and then settles, because open-ended motion is what causes Kling's notorious hang at ninety-nine percent. For image-to-video, describe only the motion and leave the scene to the image; for text-to-video, build the full scene. Lean on the negative-prompt field for the usual failure modes, morphing, warped limbs, frozen lips, floating objects. Kling is at its best on cinematic b-roll, character animation, music videos, product demos, and short social ads, and at its worst on long multi-action choreography, legible in-frame text, and anything the filters will not allow.

08

Where Kling is headed

The spin-off is the tell. Kuaishou is setting Kling up as a standalone company with outside capital and an IPO on the horizon, which means the pressure to grow revenue and defend the number-two position, at home behind ByteDance's Seedance and globally against Google, is about to get sharper, not softer. Expect the roughly quarterly release cadence to continue, and expect the next moves to target exactly the gaps above: longer clips, audio that works beyond English and Chinese and composes with the other inputs, and better physical realism, since those are the concrete places it trails.

The bet Kuaishou is making is that owning a giant short-video platform, its data on one side and its audience on the other, plus a deliberately commercial, fast, accessible posture, is enough to hold a top-tier global position even without the very best raw quality. The early evidence for that bet is the adoption: a Filmmaker Initiative launched at Cannes, an exclusive tech-partner credit on an animated feature, tens of thousands of enterprise customers, and creator contests drawing thousands of entries from over a hundred countries. Kling may not top the leaderboard, but it is one of the very few video labs that has turned a frontier model into a real, growing business, and that is its own kind of moat.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. What practical advantage did Kling have over Sora when it launched?

  2. 2. What does Kling's 3D VAE achieve with synchronous spatiotemporal compression?

  3. 3. Why did Kuaishou choose a full-attention mechanism over factored attention for temporal modeling?

  4. 4. What did Kling 2.1 introduce?

  5. 5. What makes Kling 3.0 different architecturally from earlier versions?

  6. 6. What is identified as Kuaishou's core competitive advantage in video generation?

  7. 7. What did Kling's June 2024 release demonstrate to Western observers?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.