Contents

70 / 153

Comparative Analysis: How These Models Actually Differ

The numbers, what the leaderboards actually say

Chapter 69

4 min read

Reviewed v78 · August 2026

Public benchmarks should be taken with a grain of salt, we will get to why in a moment, but they are a useful starting point for comparing models, and the patterns in them are real. The two most-watched leaderboards are the Artificial Analysis text-to-image and text-to-video benchmarks, which use human preference voting (people are shown two outputs from different models for the same prompt and asked which they prefer), aggregate the votes into ELO-style ratings, and publish the results.

Here are approximate snapshots as of early 2026, drawn from those benchmarks and from the public claims of the labs themselves. Numbers will be out of date by the time you read this, but the rough ordering is what matters.

IMAGE MODEL LEADERBOARD

Model

Lab

Architecture

Open weights

Strengths

Nano Banana Pro

Google

Gemini-native

No

Text rendering, world knowledge, 14-image composition

FLUX 1.1 [pro]

Black Forest Labs

12B Multimodal Diffusion Transformer (MM-Diffusion Transformer (DiT))

No (pro)

General photoreal, prompt adherence

Ideogram 3.0

Ideogram

Latent diffusion

No

Typography, design, posters

Recraft V3

Recraft

Latent diffusion + vector

No

Long-form text, vector output

GPT-Image (4o)

OpenAI

LLM-native

No

Instruction following, edits

Seedream 5.0

ByteDance

MM-DiT

No

Cinematic, e-commerce design

FLUX [dev]

Black Forest Labs

12B MM-DiT

Yes (non-comm)

Open ecosystem workhorse

Qwen-Image 2.0

Alibaba

7B MM-DiT

Yes

Chinese text, open license

SD 3.5 Large

Stability AI

8B MM-DiT

Yes

Permissive license

Midjourney V7

Midjourney

DiT (closed)

No

Aesthetic distinctiveness

VIDEO MODEL LEADERBOARD

Model

Lab

Audio

Max length

Cost / sec

Veo 3.1

Google

Yes (joint)

8s @ 1080p

$0.50 - $1.50

Kling 3.0

Kuaishou

Yes (joint)

15s cinematic

$0.30 - $0.80

Sora 2 Pro

OpenAI

Yes (joint)

25s @ 1792x1024

$0.50 - $2.00

Seedance 2.0

ByteDance

Yes

12s @ 1080p

$0.20 - $0.60

Hailuo 2.3

MiniMax

Partial

10s @ 1080p

$0.05 - $0.15

Runway Gen-4.5

Runway

Separate

10s @ 1080p

$0.25 - $0.60

Wan 2.5

Alibaba

Yes

10s @ 1080p

Free (open)

Luma Ray 3

Luma

Separate

10s @ 4K

$0.30 - $0.80

Pika 2.2

Pika

Separate

10s @ 1080p

$0.10 - $0.30

Hunyuan Video

Tencent

Separate

Variable

Free (open)

Notice the bifurcation. The top of the leaderboard is dominated by closed-source models from labs with massive resources (Google, Kuaishou, OpenAI, ByteDance). The middle is a mix of mid-priced commercial offerings (Hailuo, Runway, Luma, Pika). The bottom is open-source models (Wan, Hunyuan) that are free to use but require you to bring your own infrastructure. There is no single 'best' model, there is a quality-cost-openness trade-off, and your right answer depends on which axis matters most to you.

TRAINING COST COMPARISON

Model

Year

Type

Estimated cost

Compute scale

Stable Diffusion 1.5

2022

Image

~$600K

150K A100-hours

SDXL

2023

Image

$2-3M

Several hundred K A100-hours

Stable Diffusion 3

2024

Image

$5-10M

Multi-week multi-cluster

FLUX (BFL)

2024

Image

$10-30M (est)

Undisclosed

Open-Sora 2.0

2025

Video

$200K

Optimized open-source

Hunyuan Video

2024

Video

$5-15M (est)

13B params

Sora 1

2024

Video

$100M+ (est)

Months on thousands of GPUs

Veo 3

2025

Video

$100M+ (est)

Google internal

Kling 3.0

2026

Video

Undisclosed

Kuaishou internal

MODEL FUNDING & VALUATION

Lab

Total raised

Last valuation

Notable backers

OpenAI

$60B+

$500B+

Microsoft, Thrive, Khosla

Anthropic

$25B+

$170B+

Google, Amazon

Black Forest Labs

$450M+

$3.25B (Dec 2025)

a16z, Salesforce, AMP

Runway

$300M+

~$3B (est)

Google, Nvidia, Salesforce

Pika Labs

$135M

~$700M

Lightspeed, Spark, a16z

Stability AI

$100M+

Recovering

Coatue, Lightspeed

Luma Labs

$70M+

~$300M

a16z, Amplify

Ideogram

$80M+

Undisclosed

a16z, Index

Midjourney

$0 (bootstrapped)

Profitable

None - founder funded

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. How do the two most-watched leaderboards actually rank models?

  2. 2. What bifurcation does the video leaderboard reveal?

  3. 3. How did Open-Sora 2.0 train for roughly 200,000 dollars, about a thousandth of Sora 1's cost?

  4. 4. What is unusual about Midjourney's place in the funding table?

  5. 5. Why does the chapter say there is no single best model?

  6. 6. What is the catch with Open-Sora's cheap training success?

  7. 7. Which video model is listed as the cheapest paid (commercial) option per second on the leaderboard?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.