Research Deep-Dive · Explainer by Xiaohu

GPT-5.6 Sol Tops Web Design Leaderboard: It Has Design Taste and Avoids AI Giveaways

Third-party review platform Design Arena mapped 1,000 of its web pages and found that the gaps in its design space are exactly where AI clichés like purple gradients tend to live.
The TL;DR
  • OpenAI's GPT-5.6 Sol ranked #1 on Design Arena's web design (non-agentic) leaderboard, climbing 18 spots over its predecessor GPT-5.5 — the first time an OpenAI model has topped this chart.
  • Reviewers turned 1,000 of its generated web pages into image vectors and clustered them, revealing visible holes in its design map — holes that correspond exactly to common AI tells like purple gradients, bento grids, and oversized headlines.
  • Both it and GLM-5.2 avoid these clichés, but in fundamentally different ways: GLM-5.2 never learned them, while GPT-5.6 Sol learned them and deliberately refuses to use them.
  • It starts from a set of reliable templates, then heavily customizes for each brief, sitting between generic and chaotic.
  • It doesn't escape everything: confetti appears in 26.5% of outputs, and it's notably weak at data visualization.
1An Unexpected Leader

You can spot AI websites on sight. This model learned not to make them.

You can usually spot an AI-generated website from across the room: the purple-to-blue gradient background, the screen packed with bento-style cards, the giant headline dominating the fold. See enough of them, and that plastic, "clearly AI-made" look becomes almost a birthmark.

OpenAI's new GPT-5.6 Sol just took first place in web design on third-party review platform Design Arena, jumping 18 spots over its predecessor GPT-5.5 — a first for an OpenAI model. The bigger story is what reviewers found under the hood: its biggest differentiator is actively steering clear of these clichés.

Design Arena gives different models the same web design task, then has people blind-judge the finished pages. The version people pick most often ranks higher. GPT-5.6 Sol won on the "web design (non-agentic)" track, which means one-shot generation without iterative self-refinement.

Why it matters: The ranking is just the headline. Reviewers went further, converting all 1,000 outputs into image vectors and drawing a design map — a direct visual trace of a model that knows what to draw, yet chooses not to.
#1
GPT-5.6 Sol's rank on the non-agentic web design chart, 18 spots ahead of its predecessor
1,000
GPT-5.6-generated pages reviewers used to draw the design map
GPT-5.6 Sol design benchmark overview
Design Arena's benchmark overview for GPT-5.6 Sol. Source: Design Arena
2What Exactly Is Being Avoided

The giveaways that scream "AI made this"

Purple gradients and bento boxes are just two examples. Design folks call these tired conventions anti-patterns, or design smells. To understand what's new here, it helps to see the full list of smells on the table.

Three months back, Design Arena analyzed GPT-5.5 and drew up a list of these recurring tells. The main ones:

Design SmellWhat It Looks Like
Purple-blue gradientBackgrounds washed in that signature AI palette
Bento box layoutScreen divided into a grid of differently-sized cards, packed edge to edge
Oversized headlineNo hero image; just one massive line of text filling the viewport
Offset layoutElements deliberately misaligned left and right to fake a "designed" feel
Grid backgroundA faint lattice of lines layered across the entire page

Individually, none of these are inherently wrong. The problem is that AI models lean on them so heavily that readers instantly clock them. In Design Arena's blind tests, pages carrying these smells were consistently the ones humans rejected.

losing design with purple blue gradient
Smell #1: Purple-blue gradient. Pages like this routinely lost in blind judging. Source: Design Arena
losing design with grid background
Smell #2: Grid background. Source: Design Arena
3Holes in the Design Space

Turning 1,000 Web Pages into a Map — and Finding Missing Pieces

"Good taste" and "bad taste" can feel subjective. So Design Arena wanted visible proof: they plotted all 1,000 of GPT-5.6 Sol's generated pages as points on a map.

The process had two steps. First, each page screenshot was fed into a model called CLIP, which outputs a long string of numbers for every image. Think of it as a genetic code for visual style: similar-looking images get similar codes. Second, a technique called UMAP flattened those long codes into points on a 2D plane, preserving the original distances as best it could. Pages with similar styles ended up as neighboring points on the map.

An Analogy

CLIP gives each page a style ID card — a long code. UMAP then flattens that code into a 2D point, like collapsing a 3D nebula into a star chart: clusters, sparse regions, and empty voids all survive the flattening.

Laid flat, those 1,000 points form the model's design space — the full range of styles it likes to generate. And there, reviewers spotted something unexpected: several obvious voids, like chunks missing from a point cloud.

GPT-5.5's Design Map Continuous point cloud, no gaps Includes purple gradients, bento boxes, all generated normally GPT-5.6 Sol's Design Map Same spread, but with missing chunks Purple Gradient Should be here. Missing. Bento Box
Two point clouds that should look roughly the same. GPT-5.6 Sol (right) has sudden voids where points vanish; dashed circles mark the "holes." Illustrative, based on methodology described by Design Arena.

Below are the actual projections from Design Arena. The first is GPT-5.6 Sol — you can spot the gaps in the dense point field. The second is GPT-5.5, where the cloud is uniformly filled.

GPT-5.6 Sol UMAP projection with holes
GPT-5.6 Sol's UMAP projection shows obvious voids in the point cloud. Source: Design Arena
GPT-5.5 UMAP projection filled
The same projection applied to GPT-5.5: fully populated, no such gaps. Source: Design Arena
What the Holes Mean

Projections like UMAP preserve voids from the original high-dimensional space. So when one model's map has holes and another's doesn't, it means GPT-5.6 Sol can generate in that region but chooses not to.

To confirm what was in those holes, reviewers overlaid outputs from both models on a single plot: GPT-5.6 Sol's points in orange on top of GPT-5.5's points. Where orange dots appear, GPT-5.6 Sol ventures. Where only the base color remains, that's what it avoids but its predecessor visits.

overlap of GPT-5.6 Sol orange and GPT-5.5 blue dots showing purple gradient cluster gap
Orange points (GPT-5.6 Sol) overlaid on blue (GPT-5.5). The two overlap broadly, except for the purple-gradient cluster, which is completely free of orange — a region GPT-5.6 Sol entirely avoids. Source: Design Arena

The two models overlap in most regions, except for the purple-gradient cluster, where there are zero orange points. The same pattern holds for bento boxes, oversized headlines, and offset layouts. The holes are precisely where the AI tells live.

4Learned and Refused vs. Never Learned

Avoiding the Same Bad Habits by Two Very Different Routes

GPT-5.6 Sol isn't the only model that avoids these clichés. But the way it does is different — and it's the most counterintuitive finding here.

Take GLM-5.2, which also ranks high. It rarely produces oversized headlines or other tells, but its method is simple: it learned from a batch of high-scoring templates that never contained those patterns in the first place. Its design space never had a "purple gradient" region, so it can't generate one, and no hole appears on its map — because that area was always empty.

GPT-5.6 Sol is a different story. The holes on its map suggest it has learned these patterns and can draw them — it just steers around them every time. It knows what a purple gradient looks like, knows where to place it, and then decides not to.

GPT-5.6 Sol · Learned, then Refused

Its design space covers where these clichés would be, but it avoids generating there, leaving holes. It "knows, and doesn't do."

GLM-5.2 · Never Learned

It draws from a template set with no bad habits, so that style region was never in its capabilities. No corresponding zone exists on its map. It "never was an option."

GLM-5.2 avoids anti-patterns via templates
GLM-5.2 avoids bad habits through templates; that region simply isn't in its design space, so no hole appears. Source: Design Arena

Both can produce pages that dodge clichés. The difference is the shape of capability: one fenced off that land and labeled it "don't go here"; the other never mapped it at all. Design Arena suggests GPT-5.6 Sol's behavior is closer to learning followed by active suppression — a rarer trait among current models.

5Between Template and Bespoke

Starting from Reliable Templates, Then Deviating Significantly

Here's Design Arena's second finding, about personalization. Web-design models tend to fall into two camps: heavily template-dependent (stable but repetitive) or nearly template-free (varied but less reliable). GPT-5.6 Sol sits between them.

It starts with proven structures, then makes substantial adjustments for each brief. Under the same template archetype, it grows a family of related but distinct variations. Design Arena's metaphor: like bacteria evolving into closely related strains — a shared foundation, each branching its own way.

Highly Template-Based Stable, but often repetitive Almost No Templates Fully custom, highly varied GLM-5.2 Claude Fable 5 GPT-5.6 Sol Templates as base, then heavy customization
The spectrum of templating: GLM-5.2 leans toward the template end, Claude Fable 5 toward the custom end, GPT-5.6 Sol sits center-custom. Illustrative, based on Design Arena's description.

For reference points at the extremes: GLM-5.2 scored well this cycle, relying on a set of high-scoring templates; reliable, but variation mostly comes from the brief itself, and the same templates recur frequently. Claude Fable 5, by contrast, shows almost no template traces. Its design space is more dispersed, and each output is highly tailored to the brief.

GLM-5.2 templated designs look similar
GLM-5.2's outputs show visible template patterns and look similar to each other. Source: Design Arena
Claude Fable 5 varied non-templated designs
Claude Fable 5 shows almost no templating; design space is more spread out. Source: Design Arena

GPT-5.6 Sol works the middle ground: templates protect the floor, customization creates the differentiation. Reviewers believe this is a big reason for its high ranking — users get a page that fits the brief and still feels professionally crafted. One tell: when assigning images across different pages, it often reuses the same image in several different contexts.

6What It Didn't Dodge

Confetti Everywhere, Charts Still Weak

Active avoidance isn't a superpower GPT-5.6 Sol applies everywhere. Reviewers flagged two clear weak points.

First, confetti. It loves scattering confetti animations across pages — present in over 26.5% of its outputs. It even hand-rolls its own confetti library from scratch just to use it. That in itself is an AI tell, and this one it didn't avoid.

GPT-5.6 Sol overuses confetti
GPT-5.6 Sol overuses confetti, present in over a quarter of its outputs. Source: Design Arena

Second, data charts. Its performance on charts and data visualization is notably weaker — with chart.js (a common web charting library), it struggles to produce a decent real-world chart.

GPT-5.6 Sol weak at charts
GPT-5.6 Sol is weaker at charts and data visualization. Source: Design Arena
Actively Avoided

Purple gradients, bento boxes, oversized headlines, offset layouts, grid backgrounds — all left holes in the map.

Didn't Dodge

Confetti shows up in 26.5% of outputs, and it writes its own confetti library when needed; data charts remain weak.

7Faster and Cheaper

More Than Twice as Fast as the Old Leader, at Half the Price

Beyond ranking and taste, GPT-5.6 Sol also wins on speed and cost. Design Arena says it sets a new Pareto frontier on both "quality vs. speed" and "quality vs. price."

2.44×
Generation speed vs. previous leader GLM 5.2
36%
Generation speed improvement over Claude Fable 5
$5 / $30
Price per million tokens (input / output)
$10 / $50
Claude Fable 5's price for the same — 2× more expensive
What's a Pareto Frontier?

Imagine a plotted boundary: stand on it, and to make a page better you must spend more time or money — no free lunch. GPT-5.6 Sol pushed that boundary outward: same quality, faster and cheaper.

GPT-5.6 Sol pareto frontier preference vs speed and price
GPT-5.6 Sol sets a new Pareto frontier on both "quality vs. speed" and "quality vs. price." Source: Design Arena
8What This Means for Model Choice

More Discerning, and More Adaptable

Design Arena sums up its findings with two observations.

First, it's more discerning: it's learned which patterns make a page instantly read as AI-generated, and actively suppresses them — while keeping reliable structures in reserve. Second, it's more adaptable. It merges templates with customization: templates as a safety net, heavy adaptation per brief, so the result feels both tailored and professional.

Together, those two traits form the core of Design Arena's explanation for its single-round lead. As for what it hasn't dodged — the confetti, the weak charts — that's in the report too. The full picture requires both sides.

GLM-5.2 is like a model that never learned these design smells, so it can't produce them. GPT-5.6 Sol is like a model that learned them, knows they exist, and refuses to draw them. Design Arena, analysis of GPT-5.6 Sol
Source: Design Arena, "We Put GPT-5.6 Sol Through a Web Design Benchmark," July 16, 2026. All rankings, image clustering analysis, speed and price data came from Design Arena's own leaderboard and analysis; images as credited. GPT-5.6 Sol, GPT-5.5, GLM-5.2, and Claude Fable 5 are model names from their respective vendors.