Contents

85 / 153

Evaluating Quality

The dimensions of quality

Chapter 84

3 min read

Reviewed v78 · August 2026

WHY THIS SECTION EXISTS

Almost every conversation about generative imagery quality assumes there is a single quality dimension, that one model is 'better' than another in some absolute sense. This is wrong, and the wrongness causes systematic mistakes in product decisions, in model selection, and in evaluating generated outputs. Quality is multidimensional, and the dimensions that matter are different for different use cases. A model that is excellent for editorial illustration may be terrible for product photography. A model that wins benchmark comparisons may produce outputs that fail the brand-safety review of a Fortune 500 client. This section unpacks what quality actually means in this space, how to evaluate it for your specific use case, and how to think about brand-safe generation for commercial work.

A note on how this section relates to the two previous sections. The Craft of Generation covers the ten dimensions of good versus great from the perspective of the artist producing the work. The Operator's Playbook covers the infrastructure and business decisions of running a generative imagery company. This section covers the orthogonal question of how to evaluate whether a given output is actually good for a given use case, which matters for both practitioners and operators but is not reducible to either the craft or the business framing. If you are reading linearly, you arrive here with an intuition for what great work looks like (from Craft) and a plan for how to ship it at scale (from Operator), and this section gives you the vocabulary for assessing outputs against specific commercial and creative requirements.

Quality in generative imagery has at least eight distinct dimensions, and any given model is good at some and bad at others. Picking the right model for a use case is mostly a question of matching the use case's quality requirements to the model's quality profile.

01

The eight dimensions

Dimension

What it measures

Best for

Photorealism

Indistinguishable from a photo

Product, editorial, advertising

Stylistic distinctiveness

Has a recognizable aesthetic identity

Artistic, brand-driven work

Prompt adherence

Output matches text prompt accurately

Complex multi-element compositions

Compositional control

Respects layout and spatial language

Editorial layout, design

Text rendering

Generates legible, accurate text in images

Posters, signage, design

Coherence

Internally consistent (anatomy, physics, etc)

Anything with people or complex objects

Speed

Time per generation

Real-time, interactive, iteration-heavy work

Cost

Dollars per generation

High-volume production

Notice that no single model is best on all dimensions. FLUX is excellent on photorealism, prompt adherence, and coherence; weak on stylistic distinctiveness (its outputs are technically correct but aesthetically generic); medium on text rendering. Midjourney is excellent on stylistic distinctiveness; weaker on prompt adherence (it interprets prompts loosely in service of its aesthetic); medium on coherence. Ideogram is excellent on text rendering (it was specifically built for this) but weaker than FLUX on photorealism. Nano Banana Pro is excellent on prompt adherence and compositional control because it inherits Gemini's language understanding.

The right question for a use case is: which two or three of these dimensions matter most for me, and which model has the best profile on those specific dimensions? Asking 'which model is best' without specifying a use case is the wrong question.

Check your understanding

pass: 5 of 7

Answer at least 5 of 7 correctly to unlock the next chapter.

  1. 1. How many distinct dimensions of quality does the chapter say generative imagery has?

  2. 2. On which dimension is FLUX described as weak?

  3. 3. What is Midjourney's relative weakness?

  4. 4. Which model was specifically built to excel at text rendering?

  5. 5. Why is Nano Banana Pro strong on prompt adherence and compositional control?

  6. 6. The text-rendering dimension is listed as best for which use cases?

  7. 7. What is the right question to ask when choosing a model for a use case?

The weekly briefing

Get the week's moves in your inbox.

A short, sourced digest of what actually moved across generative AI, every week. Free.

Free. One email a week, no spam, unsubscribe anytime. Prefer a reader? RSS.