Skip to content

All experiments

Three briefs given to a model, then judged and rebuilt — what separates work that looks good from work that is good.
AI, 2026

Design judgment

What

A model was given three one-line briefs and produced three pieces. Each one is judged here against the same things — hierarchy, typography, spacing, interaction, accessibility — and then rebuilt. The point is not that a model designs badly. It is that "attractive" and "resolved" are different claims, and only the second can be argued.

As generated – a dashboard of four equal cards above a table and four equal buttons
1 / 6
As generated – a dashboard of four equal cards above a table and four equal buttons

01 — Dashboard

Brief: "A delivery operations dashboard for a logistics startup."

What came back. Four metric cards of identical weight, a table, four buttons of identical weight, an all-caps eyebrow, a gradient on the primary action.

Where it fails.

  • Hierarchy. The screen has one urgent number — 47 late deliveries, up 18.9% — and it sits in the fourth card at the same size as the other three. The layout says everything matters equally, which is the same as saying nothing does.
  • Colour. Every delta is green, including the two that are bad news. Colour is carrying decoration, not meaning.
  • Interaction. Four buttons, equal in weight, none of which names the thing a dispatcher opened this screen to do. "Export report" is given the gradient; it is the least urgent action on the page.
  • Type. Card labels are 14px semibold grey and the table headers are 12px uppercase with tracking — two label systems for one level of information.

Rebuilt. The late figure takes the top left at 56px and says what it means in a sentence. The three supporting numbers sit at a quarter of its size. The late row is tinted and its tag says how late. One dark button names the actual next step — assign a driver to the area that is failing — and the rest become links. The number column is tabular so the eye can compare down the column.

02 — Poster

Brief: "A poster for a three-day design festival in Seoul."

What came back. A centred cream-gradient composition, a serif display line with one word in italic terracotta, four pill-shaped tags, a rounded call-to-action.

Where it fails.

  • It is a template, not a poster. Cream background, high-contrast serif, terracotta accent: the house style of generated design. Nothing about it belongs to this festival.
  • The information is ranked backwards. A visitor needs the dates, the place and what happens. The dates are set at 15px, below a sentence that says nothing specific ("talks, workshops, and conversations about the future of design").
  • The tags are decoration. "Talks · Workshops · Exhibition · Networking" tells no one anything they could act on.
  • Everything is centred, so the composition has no structure to hold the pieces apart.

Rebuilt. The dates join the title at display size, because a poster is read at three metres and the date is what has to survive that distance. The programme becomes a definition list with real facts — how many speakers, how many seats, opening hours. The only button is the deadline, which is the one thing that expires.

03 — Pricing

Brief: "A pricing section for a SaaS product."

What came back. Three cards, a "MOST POPULAR" gradient badge, ticked lists of different lengths, a gradient call-to-action.

Where it fails.

  • Cards defeat comparison. Pricing is a comparison task, and cards force the reader to hold four items in memory while moving sideways. Different list lengths make it worse: the reader cannot tell whether an absent line is a missing feature or just unmentioned.
  • The "most popular" badge is an assertion, not information. It is styled more loudly than the prices.
  • What is not included is invisible. Nothing says Starter has no custom domain; the reader finds out later.
  • A tick is not a value. "Advanced analytics" and "Basic analytics" differ by a word and by nothing a buyer can check.

Rebuilt. One table. Rows are features, columns are plans, and an em dash says plainly where something is absent. The unit is stated where it bites — "per editor, per month". The sentence above the table answers the first question a buyer actually has: what does every plan include, and what am I paying for.

How this is judged

The same four questions, in this order, every time.

  1. What is this screen for? Name the one thing a reader came to do, then check whether the layout makes it the easiest thing to find.
  2. What does the ranking say? Size, weight and position are claims about importance. If everything is the same size, the design has no opinion.
  3. Is colour carrying meaning? Green for a figure that is bad news is a lie told in a hurry.
  4. Can this be checked? Contrast ratios, tabular figures, target sizes, focus order. Taste ends where a measurement begins.