Most generative imagery products have a unit economics structure that surprises founders the first time they actually compute it. Here is a representative breakdown for a SaaS image generation product with a $20/month consumer subscription, serving an average user who generates 200 images per month.
Cost line
Per month per user
Notes
Inference (200 images @ $0.01)
$2.00
Using a mid-tier inference platform
Storage (200 images @ 1MB each)
$0.10
Cheap until you have many users
Bandwidth (200 images served)
$0.10
Standard CDN pricing
Stripe / payment processing
$0.88
2.9% + $0.30 per transaction
Customer acquisition (amortized)
$3.00
Paid acquisition + organic ratio dependent
Customer support
$0.50
1 in 50 users contacts support
Engineering & infrastructure
$2.50
Salaries amortized over user base
TOTAL COSTS
$9.08
REVENUE
$20.00
GROSS MARGIN
$10.92
55%
That is a healthy looking business until you notice that the inference cost is the most volatile line. If your average user generates 500 images per month instead of 200, the inference line becomes $5 and your margin compresses to 30%. If the user generates 1000 images per month, the inference line becomes $10 and your margin is 5%. Some users will generate many more than 200 images, and they are usually the most engaged and most valuable users. The economics of generative imagery products are dominated by the long tail of high-volume users, and managing those users is the hardest part of running the business.
The standard responses to this problem are: tiered pricing (charge more for high-volume users), generation caps (limit how many images per month per tier), credit systems (customers pay per generation rather than per month), and degraded quality for high-volume users (free users get a cheaper model, paid users get a better one). All of these solutions have their own problems. Tiered pricing creates friction at the upgrade boundary. Caps create frustration when they are hit. Credit systems create cognitive overhead. Quality differentiation creates two-tier perception that some users find unfair.
The sustainable model that has emerged among the more mature companies is: generous monthly cap with overage pricing (Krea), or pure credit-based pricing (most B2B tools), or premium subscriptions that include a meaningful margin on top of expected usage (Midjourney's $30/month and $60/month tiers). The naive 'pay $10/month for unlimited' model does not survive contact with power users.
The cost stack for one image generation
To make the unit economics tangible, here is the breakdown of where the money goes for one FLUX dev image generation served through a typical SaaS product.
Cost component
Approximate cost
Captured by
GPU compute (raw H100 time)
$0.0015
GPU cloud (CoreWeave, Lambda, etc.)
Inference platform markup
$0.005
fal, Replicate, etc.
Model lab license fee (FLUX dev commercial)
Variable
Black Forest Labs
Application platform infra
$0.001
Your servers and CDN
Application platform team & overhead
$0.005
Your salaries amortized
Customer-facing price
$0.01 - $0.10
What you charge
Notice that the actual GPU compute is the smallest line item. Most of the money goes to operational overhead at the application layer and to the inference platform that abstracts away the GPU management. This is why inference platform pricing matters so much, the platform markup is comparable in size to your entire labor cost per generation, and reducing it directly improves margin. It is also why running your own inference becomes attractive at scale.
The unit economics of audio
Audio forces a distinction image and video never had to make, the line between speech and song, and it shows up in how each is priced. Text-to-speech is metered against input, per character or via credits that map to characters, so speech cost scales with the volume of text you push through it. Music is priced per song or, more often, as a monthly subscription with a generation cap, so it sells a fixed ceiling of outputs rather than metering usage. As of July 2026, ElevenLabs runs on a credit-per-character model across paid tiers from a few dollars a month up to enterprise, Cartesia uses the same credit logic tuned for real-time voice, and Suno sells monthly plans around eight and twenty-four dollars with a daily free allowance, with Udio offering comparable tiers. One trap for operators: API access and the consumer interface are usually priced and licensed separately, and music platforms increasingly gate commercial download behind paid accounts, so the marketing sticker rarely equals your marginal cost inside a product.
The unit economics then split by archetype. A live voice agent is dominated by minutes of speech generated in real time, so latency-optimized synthesis and concurrency set both cost and quality, and a caller who talks over the bot still bills you for the audio already synthesized. A dubbing or localization service pays per character per language, so margin is eaten by revisions and retakes, and the winning move is to charge per finished minute while your cost is per character. An audiobook or podcast pipeline is batch rather than real time, so it can use cheaper models and off-peak compute, and its economics are a one-time render amortized over unlimited plays. A music-bed feature lives or dies on the subscription cap, because if your users each generate more than the plan allows you either eat the overage or ration generations.
The unit economics of video
Video is where generative cost stops being a rounding error. The unit that matters is not the generation, it is the finished shot, and a finished shot is many generations. A working rule from productions actually doing this is about ten generations to land one usable two to three second shot. So a two minute film, cut from roughly twenty minutes of raw material, runs to about 1,200 seconds of raw generation before a single frame is kept. That number, seconds of raw generation, is the one that sets your budget.
Multiply it by a model's price per second and the spread across models is enormous. Pushing 1,200 seconds of 720p generation through an aggregator like Fal.ai, a late 2026 breakdown from the AI film school Curious Refuge puts the ceiling prices at roughly: Seedance 2.0 at $364, LTX 2.5 Pro at $204, Kling 3.0 at $202, Minimax H3 at $96, and Luma Ray 3.2 at $72. Same film, same length, a fivefold swing on model choice alone. And the floor is not $72, it is zero: LTX 2.5 run locally overnight in ComfyUI costs only electricity.
The second lever is resolution, and it is the one most people get wrong. At Seedance 2.0's 4K rate of about $1.56 per second, twenty minutes of native 4K footage for that same two minute cut runs about $1,872. Generate the identical film at 720p and upscale it afterward, where a personal Topaz license runs about $39 a month, and the all in total lands around $139 or less, for output most viewers cannot distinguish from native 4K. That is better than a tenfold saving, and the upscale, not the native render, is the sane default.
Spending a video budget well
Once you see that raw generation dominates the cost, the tactics for controlling it become obvious. A handful of habits separate a film that costs $139 from the same film that costs $1,800.
Generate at 720p, upscale last. Native 4K is the single most expensive habit in the pipeline, and it buys almost nothing a good upscaler cannot recover. Work at 720p through the whole exploration and finishing process, and upscale only the final cut. This one change is most of the tenfold saving.
Explore free, finalize paid. The hundreds of throwaway generations you burn finding a shot do not need a premium cloud model. Run a local open weights model like LTX overnight in ComfyUI for the search, and spend per second cloud money only on the takes you have already decided to keep.
Route by stakes, not by habit. A hero shot that carries the scene earns a premium model and twenty rolls. An establishing shot or a two second cutaway does not. Matching model cost to a shot's importance, rather than running everything through your best model, routinely halves a budget.
Lock identity before you scale. The re-roll tax, generating a shot again and again because the character drifts, is where budgets quietly disappear. A character LoRA or a locked reference frame established up front turns ten rolls per shot into three.
Keyframe cheap, animate dear. A still from an image model costs cents; a second of video costs dollars. Generate and select your keyframes first, cheaply, and commit to image to video only on the frames that survive. The craft habits of bookending a shot and stacking references are, among other things, a form of cost control.
Choosing your approach: budgets by use case
The right stack depends entirely on what you are making. A few common cases, from cheapest to most demanding:
Social clip or vertical short, fifteen to thirty seconds. The cheapest capable model, Luma Ray or Minimax H3, at 720p, with light upscaling or none. Tens of dollars, same day. Volume and turnaround matter more than fidelity, and the algorithm forgives a soft frame.
Two minute narrative short. 720p plus a final upscale, a mid tier model, roughly $100 to $150 in the cloud, or effectively free run locally overnight. Budget the iteration, not the render. This is the range where taste, not tooling, decides the outcome.
Pitch, previz, or mood film. Quantity over polish. Local or the cheapest cloud model, because you want fifty rough looks to find the idea, not one clean shot nobody has approved yet. Spend nothing here so you can spend on the version that gets greenlit.
Commercial or client hero spot. A premium model, Seedance or Kling, at high resolution for the two or three hero shots that sell the piece, and a budget model for coverage. Keep the iteration budget high: a client is paying, in part, for the take you would otherwise have thrown away.
Music video or stylized piece. High shot count, forgiving of imperfection. A style LoRA plus a budget model plus heavy selection beats a premium model used sparingly. The look carries the piece more than any single frame's fidelity, so buy shots, not resolution.