Journal 041
Reproducible Is Not the Same as Reliable
AIR produced deterministic, repeatable results, but checking six assessments against the websites exposed something more important: a reproducible interpretation can still be wrong.
Reproducible is not the same as reliable
12 August 2026
Yesterday, we were preparing to do something quite ambitious with AI Readiness.
We had tested 100 websites. The assessment engine was deterministic. The same evidence produced the same scores. The methodology was versioned. The tests passed.
The next idea was to expand the work to 1,000 businesses across 10 sectors and begin building a picture of how clearly business websites communicate with AI.
Today, that project is paused.
Not because the engine stopped working.
Because we realised that working consistently is not the same as being right.
The test we hadn't done enough of
AI Readiness was designed around a simple principle: use observable public evidence rather than asking an AI to make a judgement.
The engine collected website evidence, interpreted it using deterministic rules and assessed five areas:
- Who You Are
- What You Offer
- Who You Help
- Why You're Trusted
- AI Information Access
One of the attractions of this approach was reproducibility.
Run the same evidence through the same methodology and you get the same result.
We proved that repeatedly.
Then we did something much simpler.
We took six completed assessments and read the websites ourselves.
That changed things.
Language is awkward
One hotel website said:
“we've been proud to serve our local community and visitors alike”
To a person, that is a pretty clear statement about who the business serves.
Our deterministic audience rules did not recognise it.
Another website belonged to a law firm and its homepage contained the heading:
“Extra Mile Legal Service”
Our engine confidently decided that was the organisation's identity.
On another company's website, a legal disclaimer said its international umbrella entity "does not provide services to clients”.
The engine saw the relationship between an organisation, “provide” and “services” and confidently treated the sentence as evidence of what the organisation offered.
The important word was not.
Language carries context, negation, attribution, grammar, emphasis and implication.
Rules see patterns.
And sometimes the pattern is not the meaning.
The guides still stand
We started AI Readiness because we wanted to improve Wayli and understand whether the information we published was clear enough for both people and AI systems.
It has done that job.
Along the way we built a set of questions and guides that may still be useful to others.
We're resuming our work on Wayli's main:
Prepare Before You Prompt.
Ask your AI what it can establish about your business. Ask it for the evidence. Question what it inferred. Then look at your own website and decide whether anything genuinely needs making clearer.
Because the question that started AIR still matters:
What does AI say about your business?