Contents

69 / 153

Comparative Analysis: How These Models Actually Differ

What benchmarks actually measure (and what they don't)

Chapter 68

2 min read

Reviewed v78 · August 2026

You will see public benchmarks like the Artificial Analysis text-to-image leaderboard and various video benchmarks ranking these models against each other. It is worth understanding what these benchmarks measure and what they miss.

Most public benchmarks are based on human preference, people are shown two outputs from different models and asked which one they prefer for the same prompt, and the winning rates are aggregated into ELO-style ratings. This measures something real but it has well-known biases. People tend to prefer images that look more polished, more colorful, more conventionally attractive, which means models tuned to produce 'pretty' outputs do well even if they are bad at following specific instructions or rendering accurate details. Models with strong opinions (Midjourney, Krea-1) sometimes get penalized because the AI judges or human raters do not appreciate their aesthetic choices.

Fig.diagram
ONE PROMPT“a cat ridinga bicycle”OUTPUT Amodel hiddenOUTPUT Bmodel hiddenhumanpicks Ax 1000sELO LEADERBOARD1. model A12742. model D12513. model B12404. model C1198average preference,not fitness for your jobA leaderboard measures who wins the average blind vote. It does not measure who wins on the prompt you actually need.
How preference leaderboards actually work. Two models answer the same prompt, a human picks the winner blind, and thousands of these pairwise votes resolve into an ELO score. It measures average preference, not fitness for your specific job.

More specific benchmarks, text rendering accuracy, prompt following, anatomical correctness, physical plausibility for video, give you more useful information for specific tasks but are less commonly published. The honest truth is that the best way to evaluate models is to test them on your own use cases with your own prompts and judge the results yourself. Benchmarks are useful as a sanity check and as a starting point, but they should not be the final word.

The deeper problem is that the thing most benchmarks measure, average preference across a broad prompt set, is not the thing most production work needs, which is reliability on a narrow one. A model that wins the arena by being pleasing on landscapes and portraits can still fail every time you ask for a hand holding a specific object, legible text on a package, or the same character twice. Aggregate scores hide these tail failures, and the tail is exactly where professional work lives. Two other distortions are worth naming. Benchmarks reward the median output, but a skilled operator cares about the best output out of ten, which is a different quantity entirely. And blind-preference tests strip away the controllability, editing, and workflow integration that decide whether a model is usable at all, so a model can top the leaderboard and still be the wrong tool for a real job. Read the numbers as a rough filter, not a ranking, and trust a weekend of your own prompts over any leaderboard.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. What are most public image benchmarks based on?

  2. 2. What bias affects human-preference benchmarks?

  3. 3. Why do opinionated models like Midjourney or Krea-1 sometimes score lower?

  4. 4. What is true of more specific benchmarks like text-rendering accuracy or prompt following?

  5. 5. What does the chapter say is the best way to evaluate models?

  6. 6. How should public benchmarks be treated?

  7. 7. How are the ELO-style ratings in these benchmarks produced?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.