Contents

102 / 153

The Ecosystem: A Field Guide to the Whole Board

Inside fal: one platform up close

Chapter 101

10 min read

Reviewed v78 · August 2026

fal calls itself the generative media platform for developers, and the emphasis on both words matters. Media, not language: it does not host chatbots, it hosts the image, video, audio, and 3D models this book is about. And developers: you do not log into a nice web app, you call an API, and fal's whole pitch is that it serves those models faster than anyone, on an inference engine it claims runs up to ten times quicker than a naive deployment, across thousands of H100-class GPUs that scale from zero to a fleet and back. The business has grown into that pitch: revenue up sixtyfold in a year, more than a million and a half developers, and a hundred-and-forty-million-dollar round led by Sequoia in late 2025 that valued the company around four and a half billion dollars.

What makes fal worth a chapter is not the plumbing, though, it is the catalog. Over a thousand models sit behind that one API, and because fal is a launch partner for most of the labs rather than an exclusive home for any, the catalog is close to a complete census of the field. It is where you discover that the medium is far wider than the dozen model names that make the news. A useful tell: almost everything is here, with one conspicuous exception, Runway, which keeps its models on its own platform. Everyone else rents fal's speed. Two of the models in the garden are fal's own, the open AuraFlow image model and the AuraSR upscaler, but the rest is the whole industry, shelved in one place. Walk the shelves.

01

The image bench

Start with still images, where the interesting story is not the frontier names but the specialists. Bria is the model enterprises reach for when the lawyers are watching: it is trained only on licensed data, from partners like Getty and Alamy, ships with full intellectual-property indemnification, and carries a patented attribution engine that traces each output back to the training images that shaped it and pays those owners, a Spotify model for training data. Reve, a Palo Alto lab founded by researchers out of Adobe and Google, took the opposite of the brute-force route, treating an image as something to lay out like code, planning a structured, editable composition before rendering; its model debuted anonymously under the codename Halfmoon, quietly took the number-one spot on the image arena, and by mid-2026 was a genuine frontier contender with unusually precise text and region-level editing.

Recraft, which briefly held the arena's top spot under the codename red_panda, is the designer's model: it is the one that outputs real vector art and SVGs, logos and icons, not just pixels. Ideogram remains the typography specialist. And the most quietly radical entries are about efficiency: NVIDIA's Sana compresses the image so aggressively it can render a four-thousand-pixel picture in under a second on a laptop GPU, and HiDream's open, sparse-transformer model reaches frontier quality in a handful of steps. There is even a small movement of licensed-data open models here, F Lite, built jointly by Freepik and fal on eighty million licensed images, sitting next to fal's own open AuraFlow. The lesson of the bench is that the question 'which image model' has stopped having one answer. It has a dozen, each best at one thing.

The frontier specialists keep arriving on the same shelf. Reve, the fast-rising lab, puts its Reve 2.1 here for prompt adherence and clean text rendering, and Krea ships its speed-tuned Krea 2 Turbo, which generates a high-fidelity image in seconds. The point is not that these are obscure, it is that on fal the newest architectures show up as just another address to call, next to the workhorses.

02

Editing, matting, and relighting

The edit shelf is where a lot of the real production work happens, and it is deeper than the generators. FLUX.1 Kontext, from Black Forest Labs, is the instruction editor people reach for because it holds a character and a style steady across many successive edits instead of redrawing the whole frame each time. Google's Nano Banana does conversational, multi-turn editing and blends several images into one. Alibaba's open Qwen-Image-Edit is uncanny at editing the text inside an image while keeping the original font. ByteDance's Seedream folds generation and editing into one model with up-to-six-image reference editing at 4K, and StepFun's Step1X-Edit can reason over an abstract instruction rather than a literal one.

Underneath the generators sit the quiet workhorses that are really academic research. BiRefNet, a university segmentation model, does the clean background cutout that every product photo needs, and Bria wraps that same architecture in licensed data for commercial-safe matting. IC-Light, from the researcher who also wrote ControlNet, does the thing that looks like magic in a demo and saves a shoot in practice: it relights a subject in a physically consistent way, changing the light while keeping the object intact. And a whole corner is virtual try-on, from the lightweight academic CatVTON to FASHN's production model, because putting clothes on a person convincingly is its own hard, valuable problem.

The matting and conditioning shelf goes much deeper than one segmentation model. Bria RMBG 2.0 removes backgrounds using a model trained only on licensed data, the commercially safe cutout an enterprise can actually ship, while Meta’s Segment Anything 2 auto-masks any object across an image or a video. For conditioning, Depth Anything and Marigold turn a flat frame into the dense depth map a ControlNet steers from. And the upscaler rack is a category unto itself: fal’s own AuraSR, an open reproduction of the GigaGAN upscaler, sits beside Clarity, CCSR, DRCT, and the classic ESRGAN, each a different bet on how to invent detail that was never captured.

03

The video shelf

Video is the fastest-moving shelf and the one where the open models matter most. The frontier closed engines are all here, Kling in a deep ladder of versions with real physical motion, Google's Veo, which launched its API on fal first, with native synchronized audio, ByteDance's Seedance, and MiniMax's Hailuo. But the interesting weight is on the open pole. Alibaba's Wan is the open-weight video family that actually runs on a single consumer card, using a mixture-of-experts trick to split the denoising between a high-noise and a low-noise expert. And Lightricks' LTX-2 is the standout of the open camp: a fully open model that generates synchronized audio and video together, in one pass, at native 4K, and still runs on a consumer GPU, a combination nobody else has opened up. Lightricks, the company behind the consumer apps Facetune and Videoleap, runs the open model and a polished filmmaking product, LTX Studio, off the same technology.

The single most futuristic thing on the shelf is Decart's Lucy, which is not batch generation at all: it transforms a live video stream in real time, at around thirty frames a second with almost no latency, swapping characters and restyling footage as it plays. Decart, valued around four billion dollars and counting Andrej Karpathy among its backers, builds real-time world models; its playable, engine-free generated game Oasis is the other famous one. It points at a future where generation is not a render you wait for but a runtime you act inside, which is a genuinely different thing from everything else in the garden.

The video shelf is also where restoration lives. ByteDance’s SeedVR2 restores and upscales footage while preserving natural texture, Topaz brings its long-trusted video sharpening to the API, and FlashVSR and SIMA cover the fast and the gentle ends of the same job. On the generation side the consumer engines are here too, PixVerse for lifelike physics and Lightricks’ LTX for extension and aspect-ratio reframing, so a clip can be generated, extended, and finished without ever leaving the platform.

04

The sound stage

The audio shelf is the one most readers will not have explored, and it is full. The newest arrival captures the theme: Sonilo generates a soundtrack directly from your video, with no prompt at all, reading the footage's pacing and emotional arc and scoring it, and because it is trained on professionally licensed catalogs the result is cleared for commercial use, the audio analogue of Bria's licensed-data pitch. ElevenLabs brings its whole suite, text-to-speech, sound effects, licensed music, transcription, and dubbing. Sony's MMAudio does the reverse problem, generating a synchronized soundtrack from a silent video with tight audio-visual sync. And Stability's Stable Audio makes structured, minutes-long music.

Then there is the long tail of specialists, each interesting for one reason. F5-TTS clones a voice from a few seconds of audio with no alignment step. Sonauto and CassetteAI race to generate full songs with lyrics in seconds. Google's Lyria and the open ACE-Step and DiffRhythm generate complete tracks. Nari Labs' Dia does multi-speaker dialogue with laughs and sighs written in. And tiny Kokoro does competent speech at eighty-two million parameters for about two cents per thousand characters. It is a reminder that generative AI quietly became generative audio too, and almost nobody outside the field noticed.

The voice rack alone is a tour of the open text-to-speech world. Kokoro, at eighty-two million parameters, runs many times faster than real time; Resemble’s Chatterbox clones a voice and starts speaking in under 150 milliseconds; Nari Labs’ Dia, Alibaba’s Qwen3-TTS, and MiniMax’s Speech-02 each clone zero-shot from a single reference sample. Beyond speech, Meta’s Demucs splits a finished track back into vocals, drums, and bass, Stability’s Stable Audio 2.5 generates music minutes long, and CassetteAI turns a prompt into a royalty-free stereo track in seconds.

05

Dimension, detail, and faces

Three more shelves round out the garden. Three-dimensional generation has gotten quietly good: Microsoft's TRELLIS turns a single image into a full 3D asset with real materials, Tencent's Hunyuan3D and Deemos' Rodin output clean, production-ready meshes, and Meshy and Tripo cover the game-asset pipeline including retopology. The detail shelf is the unglamorous one that makes everything else usable: fal's own AuraSR is a fast upscaler, the academic SUPIR restores a degraded photo using a diffusion prior and a text prompt, ByteDance's SeedVR2 does the same for video in a single pass, and the open Clarity upscaler is the free answer to Magnific.

And the face shelf is where a lot of AI film actually gets finished. LivePortrait animates a still portrait from a driving video using implicit keypoints rather than diffusion, so it is fast and controllable, and it works on animals as well as people. Sync Labs, ByteDance's LatentSync, and Tencent's real-time MuseTalk handle the lip-sync that marries a generated voice to a generated face. None of these are household names. All of them are load-bearing, and every one of them is a single API call away on the same platform as the frontier engines.

The 3D shelf has matured past single meshes: Microsoft’s TRELLIS now accepts LoRA adapters, and Tencent’s Hunyuan3D outputs meshes with PBR materials, retopology, and part-splitting, close to something a game team could actually drop into an engine. The face shelf is deep too, because it is where AI film gets finished: Sync Labs’ Lipsync 2.0 and VEED’s lipsync match any audio to any face, Tencent’s MuseTalk does it in real time at thirty frames a second, and SadTalker animates a single portrait into a talking head from nothing but a voice track.

06

How you actually use the garden

The mechanics are worth knowing because they are the point. Every model has an address, a namespace and a path like fal-ai/flux/dev or fal-ai/bytedance/seedance, and you call it through an async queue that holds your request, retries it if a GPU hiccups, and hands back the result by poll or webhook, which is what lets an app lean on a thousand models without operating a single GPU. You can train your own LoRA on fal in minutes and get a private model back. You can chain models into a workflow so one endpoint runs a whole pipeline, generate, then upscale, then relight, then lipsync, passing each output to the next. And you pay two ways: by unit of output for the hosted models, a fraction of a cent per megapixel of image or a dime or two per second of video, or by the GPU-hour if you bring your own model to fal's engine.

The reason the garden matters, in the end, is the reason the stack chapter kept insisting on: the model is a commodity, and the value is in knowing the board. fal is the board made legible. A working operator who has walked these shelves knows there is a licensed-data model for the nervous client, a real-time model for the live demo, a vector model for the logo, a relighting model for the badly lit shot, and a video-to-music model for the edit, and that the difference between an amateur and a professional is increasingly just knowing which door to open. The frontier gets the headlines. The garden is where the work gets done.

Check yourself0 / 4

Q01

What makes fal worth studying beyond its speed?

Q02

What does Runway's absence from fal signal?

Q03

Why has the question of which image model to use stopped having one answer?

Q04

How do you actually use the fal garden mechanically?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.