Reproducible Is Not the Same as Reliable
AIR produced deterministic, repeatable results, but checking six assessments against the websites exposed something more important: a reproducible interpretation can still be wrong.
For the practical website journey behind the research, see the historical Wayli case study, or read the AIR methodology historical reference.
AIR produced deterministic, repeatable results, but checking six assessments against the websites exposed something more important: a reproducible interpretation can still be wrong.
Making AIR's scoring more transparent exposed an Audience score that could never be reached. Testing the problem led us to simplify the question rather than weaken the evidence standard.
Testing AIR on Wayli exposed a weakness in where the assessment looked, leading to a more representative bounded acquisition method.
A technically healthy website can still produce a poor AI answer when the retrieval system selects the wrong evidence.
As AI becomes a new audience for the web, websites are beginning another period of adaptation. This is why we're choosing to build AI Readiness in public.
Six independent AI reviews reached different scores but repeatedly identified the same questions about clarity, audience, trust and evidence.
An independent review of Wayli by ChatGPT, assessing the platform from the perspective of how modern AI systems evaluate and interpret websites.
Claude's independent review of Wayli's crawlability, content, trust evidence and citation viability.
Perplexity's independent review of Wayli's positioning, design, content depth, navigation and trust foundations.
Gemini's independent review of Wayli's human value, AI processability, interoperability and external authority.
GitHub Copilot's review of Wayli's user experience, repository architecture, AI context model and engineering priorities.
Microsoft Copilot's cross-tab review of Wayli, followed by a scored assessment of the evidence visible on the AI Readiness page.
While testing Wayli with browser AI, we discovered something unexpected: sometimes AI isn't looking at the same information as you.
We discovered our own assessment could never award a perfect score. The problem wasn't the methodology. It was our implementation.
Building an AI Readiness product taught us something unexpected: the most trustworthy AI report might be the one that AI never writes.
The biggest improvement to AI Readiness wasn't a better engine. It was realising what people should not have to pay for.
The biggest improvement to our latest AI Readiness report wasn't caused by changing the engine. It happened because we acted on what the previous report told us.
We thought we were improving an AI Readiness score. What we were really improving was how clearly we explained ourselves.
A plain English explanation of how AI forms an understanding of a business and how Wayli measures it.
Testing Wayli didn't just improve our website. It changed what AI Readiness was becoming.
Testing one website taught us about Wayli. Testing one hundred began to teach us about AI.
Our first benchmark suggests many businesses may be harder for AI to understand than their owners realise. It's an early observation—not yet a conclusion.
We thought we'd find one big problem. Instead, we found lots of small ones. That changed how we think about AI Readiness.
One hundred businesses gave us our first evidence. It also reminded us how much we still have to learn.
Before asking anyone else to trust our Understanding Engine, we pointed it at ourselves.
We set out to improve a mortgage page and ended up questioning whether Wayli should generate reports at all.
The report looked incomplete, but the missing understanding was already there.
The biggest breakthrough wasn't teaching AI to make better decisions. It was deciding that it shouldn't make them at all.
AI Readiness wasn't built by looking for customers. It started by trying to understand whether AI understood Wayli.
We weren't trying to build another product. We were trying to understand how AI would understand Wayli.