Journal 041
Reproducible Is Not the Same as Reliable
AIR produced deterministic, repeatable results, but checking six assessments against the websites exposed something more important: a reproducible interpretation can still be wrong.
Journal entries connected by topic.
Journal 041
AIR produced deterministic, repeatable results, but checking six assessments against the websites exposed something more important: a reproducible interpretation can still be wrong.
Journal 040
Making AIR's scoring more transparent exposed an Audience score that could never be reached. Testing the problem led us to simplify the question rather than weaken the evidence standard.
Journal 039
Testing AIR on Wayli exposed a weakness in where the assessment looked, leading to a more representative bounded acquisition method.
Journal 029
As AI becomes a new audience for the web, websites are beginning another period of adaptation. This is why we're choosing to build AI Readiness in public.
Journal 030
Six independent AI reviews reached different scores but repeatedly identified the same questions about clarity, audience, trust and evidence.
Journal 031
An independent review of Wayli by ChatGPT, assessing the platform from the perspective of how modern AI systems evaluate and interpret websites.
Journal 032
Claude's independent review of Wayli's crawlability, content, trust evidence and citation viability.
Journal 033
Perplexity's independent review of Wayli's positioning, design, content depth, navigation and trust foundations.
Journal 034
Gemini's independent review of Wayli's human value, AI processability, interoperability and external authority.
Journal 035
GitHub Copilot's review of Wayli's user experience, repository architecture, AI context model and engineering priorities.
Journal 036
Microsoft Copilot's cross-tab review of Wayli, followed by a scored assessment of the evidence visible on the AI Readiness page.
Journal 027
While testing Wayli with browser AI, we discovered something unexpected: sometimes AI isn't looking at the same information as you.
Journal 028
We discovered our own assessment could never award a perfect score. The problem wasn't the methodology. It was our implementation.
Journal 026
Building an AI Readiness product taught us something unexpected: the most trustworthy AI report might be the one that AI never writes.
Journal 024
The biggest improvement to our latest AI Readiness report wasn't caused by changing the engine. It happened because we acted on what the previous report told us.
Journal 023
We thought we were improving an AI Readiness score. What we were really improving was how clearly we explained ourselves.
Journal 019
Testing one website taught us about Wayli. Testing one hundred began to teach us about AI.
Journal 020
Our first benchmark suggests many businesses may be harder for AI to understand than their owners realise. It's an early observation—not yet a conclusion.
Journal 021
We thought we'd find one big problem. Instead, we found lots of small ones. That changed how we think about AI Readiness.
Journal 022
One hundred businesses gave us our first evidence. It also reminded us how much we still have to learn.