In my last piece, I looked at what changes when teams take pre-launch testing seriously: Earlier sign-off. Clearer ownership. Cleaner experiments.
That article was really about workflow. This one is about something bigger.
AI has dramatically lowered the cost of creating options. The challenge is no longer about whether we can generate enough ideas. The challenge now is how do we decide which ones are worth pursuing
And that raises a question I don't think the industry has fully answered: who - or what - is doing the judging?
The default answer is a bad one
Right now, for most teams, the answer is simple: whoever happens to be in the room. Generative AI can produce fifty variants of an ad, a landing page, a headline, in the time it used to take to brief one. Someone still has to decide which of those variants is worth shipping. Increasingly, that decision comes down to a gut reaction from whoever has the strongest opinion or, perhaps, the most senior title in the room.
This isn’t a rigorous evaluation system, it's a new bottleneck. AI hasn't removed the constraint but has moved it. The problem is no longer generating enough ideas but deciding, consistently and objectively, which ones to move forward.
Replacing one bottleneck with another isn't progress - it just means we're producing more ideas than our decision-making process can comfortably handle.
Generation and Evaluation are different jobs
The same system shouldn't both generate ideas and decide which ones are best.
Here's the central argument of this piece: the same system shouldn't both generate ideas and decide which ones are best.
Not because one is "better" than the other but because they're solving fundamentally different problems. A generative model is trained to create outputs that resemble the examples it has learned from. Its job is to generate plausible, useful creatives.
An evaluation system has a different job entirely. It predicts how real people are likely to respond to a specific piece of creative. Will they notice it? Understand it? Remember it? Act on it?
Those questions can't be answered by generation alone - they require behavioural evidence and models trained on human responses.
Being excellent at generating creative doesn't automatically make a system good at evaluating it. It's not a feature that's missing or a capability that simply needs switching on. It's a different discipline, built on different data and designed to answer a different question.
This isn't a new problem. Creating content is something the biggest platforms in the world solved for themselves years ago, at a scale most of us aren't thinking about.
The platforms solved this years ago
The world’s biggest digital platforms have been separating content from evaluation for years.
Take Meta. Creators and brands produce the posts and videos that enter its ecosystem. But Meta uses separate ranking systems to decide what gets surfaced to each user. Those systems learn from behavioural signals such as watch time, likes and shares to predict what people are most likely to find relevant.
TikTok works on a similar principle. Creators make the videos; its recommendation system evaluates and ranks them using signals including likes, shares, comments and whether someone watches a video through to completion.
In both cases, creation and evaluation are distinct jobs. One supplies the content; another uses evidence of human behaviour to decide what deserves attention.
That separation isn’t an implementation detail but is fundamental to how these platforms operate at scale. They don’t rely on the content creation process itself to determine what deserves attention. They have a separate evaluation layer, informed by how real people behave.
The brand side of the industry - marketers, agencies and in-house creative teams - is only now beginning to face a similar challenge. As generative AI produces more creative options than ever before, teams need an independent way to evaluate those options with the same kind of rigour.
The difference is that most brands don’t have billions of live interactions generating behavioural signals every day. They need a way to evaluate creative before those signals exist.
Where EyeQuant fits

So where does EyeQuant fit into all of this? We think of EyeQuant as the evaluation layer for teams that don't have a billion daily users generating behavioral signals for them.
The biggest platforms generate behavioural data as a by-product of their scale. Every click, pause and scroll becomes another signal they can use to improve their ranking systems.
Most brands running a generative AI creative pipeline don't have that luxury. They have creative teams producing more options than ever before, but no equivalent evaluation layer to help them understand which ideas are most likely to work before they go live.
That's the problem EyeQuant was built to solve. Over more than a decade, we've developed models of human visual attention trained on behavioural and neuroscience data. We predict how people are likely to respond to a creative: what they'll notice first, whether the visual hierarchy works and whether the design supports the intended outcome.
In other words, it gives creative teams something most organisations don't have: an independent way to evaluate work before it reaches a live audience.
Why not just add evaluation to the generator?
It's a reasonable question. If generative AI is improving so quickly, couldn't the same company simply add an evaluation feature? Yes, most probably. Adding a button isn't the hard part. The hard part is building an evaluation system grounded in real human behaviour rather than the generator's own output.
Generation models learn patterns in content. Evaluation models learn patterns in human attention and behaviour. Those are different problems, trained on different data. That's why recognising the need for better judgement isn't the same as delivering it.
Canva's recent State of Marketing and AI 2026 report makes exactly this observation, describing AI as creating "volume without vision." I think that's a useful diagnosis.
But identifying the problem isn't the same as solving it. Solving it requires something different: independent behavioural evidence that can predict how people are actually likely to respond.
And even if a generator could score its own work, I'd still question the architecture. The system producing the creative would still be the system judging the creative.
Whether you're reviewing your own writing, analysing your own experiment or evaluating your own design, separating creation from evaluation almost always leads to better decisions.Why should AI be any different?
Whether you're reviewing your own writing, analysing your own experiment or evaluating your own design, separating creation from evaluation almost always leads to better decisions.Why should AI be any different?
The next competitive advantage
Every major shift in technology changes where value sits. AI has made creating dramatically easier. That means the scarce resource is no longer ideas, it's judgement.
For years, competitive marketing advantage came from producing more creative, more quickly.
Increasingly, I think it will come from something else entirely: the ability to evaluate ideas independently, consistently and before they reach the market.
That's why I believe evaluative AI will become its own category, sitting alongside generative AI rather than inside it.
I think the organizations that benefit most from AI won't be the ones generating the most content. They'll be the ones best at deciding what deserves to be seen.
Next: Beyond Attention - The Behavioral Stack
Attention is where prediction starts. Next, we’ll explore what comes after it, from comprehension and engagement to action - and what it really takes to predict human behaviour with confidence.

