Production is the loud center of the field, the phase this book spends most of its pages inside, and it is far larger than the models. Start with the engines themselves. Image generation is a frontier race between Midjourney's house aesthetic, OpenAI's reasoning-driven GPT-Image, Google's Nano Banana line, Black Forest Labs' open-weight-friendly FLUX, and fast-rising labs like Reve, Ideogram, Recraft, and Krea, with Stability's Stable Diffusion still the open foundation underneath much of the ecosystem and newer open entrants like HiDream and Alibaba's Qwen-Image crowding in. Video is a parallel race between Google's Veo and its new Gemini Omni line, Kuaishou's Kling, Runway, MiniMax's Hailuo, Alibaba's open-weight Wan, and ByteDance's Seedance, with OpenAI having exited consumer video entirely. As of mid-2026 the public arenas put Google's fast Gemini Omni Flash on top of the video board, a reminder that the ranking turns over every few weeks. These are the instruments the model chapters profile in depth. On the map they are two boxes among dozens.
Around the engines sits the layer where creators actually work. Because nobody drives a raw model anymore, the interface has climbed up to the aggregators and canvases: Krea, OpenArt, Leonardo, Higgsfield, and Magnific each wrap many models behind one surface with their own defaults and taste. For power users the node-based canvases, ComfyUI foremost, with Flora and Weavy alongside, make every step of a pipeline explicit and composable. And because a raw clip is not a scene, a whole category exists just to hold a character's face and wardrobe across shots: the multi-shot and narrative tools like Katalist, Pippit, and LongStories.ai that turn single generations into sequences. Realtime is its own frontier, where Krea's live canvas and Decart's streaming models generate as fast as you can move.
The shape to notice is that the engines are surrounded on every side by businesses that exist because the engines alone are not enough.
The engines, profiled elsewhere
The image and video model labs are the engines of this whole field, and the book profiles each of them in its own chapter: Midjourney, FLUX, Nano Banana, GPT-Image, Ideogram, and the rest of the image frontier, alongside Sora, Veo, Kling, Runway, Seedance, Wan, Hailuo, and the video labs that followed. This tour covers everything built around those engines, the layer of tools that turns a raw model endpoint into something a working creator can actually use.
Generation you can steer live
Most generative video is a batch process. You write a prompt, wait, and judge the result after the fact. Realtime video collapses that loop until the model responds while you type, paint, or move, which changes the medium from a slot machine into an instrument. The category exists because latency is the difference between describing an image and performing one, and because interactive frame rates are the precondition for live tools, games, and anything a person expects to feel responsive under their hands.
Krea Realtime is the most creator-facing expression of this idea, a latency-optimized surface where the canvas updates as you paint or edit, so the generation feels like direct manipulation rather than a request. Decart pushes the frontier on the infrastructure side, running streaming diffusion at sub-40ms per frame, fast enough to drive interactive and world-model experiences like its Oasis and Mirage projects rather than pre-rendered clips.
The other two work the plumbing that makes realtime practical at scale. Livepeer brings decentralized video infrastructure to the problem, distributing the compute for realtime AI across a network rather than a single provider. Keyframe Labs sits further upstream in research, working on the interactive generation techniques that the rest of the category depends on. Together these four are betting that the next interfaces are performed, not prompted.
Where creators actually work
No serious creator lives inside a single model. The frontier moves weekly, each engine is best at something different, and switching between a dozen provider websites is its own kind of tax. Aggregators exist to be the one surface where all of that converges: a canvas or dashboard that routes to many image and video models, remembers your assets, and adds the controls the raw APIs never bother to ship. The bet is that the durable value is in the interface and the workflow, not in owning the weights.
Krea is the fullest version of this, a unified canvas spanning many image and video models plus its own foundation model, so the aggregator and the lab are the same company. Leonardo.Ai took the creator-platform route with fine-tunes and a large community, and is now part of Canva, which pulls generative tooling toward the mass design market. Higgsfield has grown fast as a multi-model video and image platform built around cinematic controls, the camera moves and shot grammar that filmmakers reach for instinctively.
The rest specialize. Magnific, now part of Freepik, pairs a multi-model canvas with best-in-class upscaling and enhancement, the finishing pass that makes generated frames hold up at size. Scenario narrows all the way to game assets, training custom models so a studio can generate in its own consistent house style. OpenArt runs a broad multi-model image playground with a community layer, the low-friction on-ramp where a lot of people first learn the medium before graduating to heavier tools.
Nodes, graphs, and the power-user stack
The infinite canvas is what generation looks like once you stop treating it as a chat box and start treating it as a pipeline. Instead of one prompt and one output, you get a spatial workspace of nodes and connections, where the output of one model feeds the input of the next, and the whole graph becomes a reusable, inspectable recipe. The category exists for the people who outgrew simple interfaces and need to see and control every step, chaining image, video, and text models into something repeatable.
ComfyUI is the anchor and, as the book argues at length in its own chapter, the single most important piece of infrastructure in the open ecosystem. Its node-based graph became the power-user standard because it exposes the full machinery of a generation and lets a community wire up workflows no product manager would have shipped. Everything else in this category is, in part, a response to how capable and how intimidating ComfyUI is.
Flora and Weavy take the node-and-canvas idea and aim it at professional pipelines with a friendlier surface, Flora chaining image, video, and text models on a single workspace and Weavy targeting studio-grade production flows. Visual Electric goes the other direction, a canvas-first image workspace built for designers who want the spatial freedom without the graph complexity. Magnific appears here too, adding its upscaling strength to a canvas, a reminder that these categories overlap wherever a tool is good at more than one job.
Holding a story across shots
A single generated clip is a moment. A story is many moments that have to agree with each other, the same character with the same face and wardrobe across a dozen shots, cut together in an order that reads as narrative. That consistency is exactly what raw video models are worst at, since each generation starts fresh. This category exists to sit on top of the engines and impose continuity, memory, and shot structure, turning a pile of clips into something that feels authored.
Katalist is a clear example, built around character-consistent multi-shot storyboarding and video so a figure stays recognizable from setup to payoff. Pippit, ByteDance's multi-shot content generator under the CapCut umbrella, brings that same shot-sequencing logic to the enormous consumer-content audience its parent already owns. LongStories.ai pushes toward the hardest end of the problem, long-form multi-shot generation where continuity has to survive across many minutes rather than seconds.
The edges of the category are more experimental. glif takes a composable approach, letting people assemble AI mini-apps and workflows that can be strung into narrative pipelines rather than shipping one fixed tool. Utopai positions itself as a full AI studio for narrative production, the ambition being an end-to-end path from story to finished piece. What unites all of them is the recognition that the interesting problem is no longer generating a good shot, it is generating the next shot that belongs with the last one.
Higgsfield: cinematic control and the aggregator bet
Higgsfield is the sharpest example of a company betting that in generative video, control and distribution beat owning the best model. It is a San Francisco creative suite founded in October 2023 by Alex Mashrabov, previously head of generative AI at Snap, whose earlier deepfake and avatar startup AI Factory was bought by Snap for a reported one hundred sixty-six million dollars. The defining product idea is directability. Instead of asking you to wrestle a raw text-to-video model, Higgsfield layers film-grammar camera moves on top of generation, a library of dozens of motion presets like crash zoom, dolly zoom, crane, FPV drone, and bullet time, so a prosumer can direct a shot like a cinematographer without owning a rig. Its own models, an image-to-video model called DoP and an image model called Soul with an identity-consistency feature, sit alongside a swappable roster of wrapped frontier engines including Kling, Veo, Sora, Hailuo, Wan, and Seedance, presented as one tool with one credit balance.
The growth curve is the part that made the industry look. Reported revenue ran from roughly fifty-eight million dollars across 2025 to a two hundred million run rate in December, to a three hundred million annualized rate weeks later, to around five hundred million by mid-2026, one of the faster consumer-AI ramps anyone has posted and enough to draw move-over-Cursor comparisons. Funding kept pace: a Series A extended by an eighty million dollar raise in January 2026 valued the company near 1.3 billion, with talks of a much larger round reported later in the year, and Mashrabov has publicly targeted a one billion dollar annual run rate by year-end. The company claims scale to match, on the order of fifteen million users and millions of video generations a day, monetized through tiered subscriptions and a credit system, with a creator-payment program and an ad-creative business layered on top.
Higgsfield, the honest read
The strength is real and specific. Getting deliberate, purposeful camera motion out of generative video is genuinely hard, and Higgsfield's preset-and-template workflow gets a non-expert to a usable, on-trend short-form clip faster than almost anything else, which is exactly why AI-influencer accounts, music videos, and high-volume ad creators gravitate to it. The weakness is structural and follows directly from the bet. Because Higgsfield mostly wraps other companies' base models, its quality ceiling largely tracks Kling, Veo, and Sora rather than exceeding them, and any pricing or policy change upstream ripples straight through its product. It is a video factory, strong at the shot, thin on the surrounding needs like music, voiceover, captions, and lipsync that a finished piece requires.
On access, Higgsfield runs as web and mobile apps on a credit-metered model, with reported tiers in the rough range of fifteen dollars a month at the low end up to around one hundred for the heaviest plan, though the exact numbers move with promotions and should be checked live. The way to get value is the same as the way to use it well: lock a consistent character or look with Soul first, then feed that still into the image-to-video path and choose one strong camera move that serves the beat rather than stacking many, and match the wrapped base model to the job. The open question for the company is whether it can climb from aggregator to owning more of its own frontier-quality generation, and whether it can repair enough trust to win the enterprise and agency spend its revenue ambitions require.