Journal 039
We Changed Where AIR Looks, Not What It Believes
Testing AIR on Wayli exposed a weakness in where the assessment looked, leading to a more representative bounded acquisition method.
AIR's stricter evidence methodology gave Wayli a fresh score of 80/100.
The five areas looked like this:
- Identity: 100
- Offer: 100
- Audience: 0
- Authority: 100
- Information Access: 100
The weak point was clear. AIR could not establish who Wayli was intended to help.
We treated that as a useful challenge to the website.
That turned out to be only part of the story.
We started by changing Wayli
First, we made Wayli's audience explanation clearer for human visitors. We distinguished between people using Wayli to explore important personal decisions and businesses using AI Readiness to examine their public website evidence.
Then we ran the same frozen assessment again.
The score remained 80. Audience remained 0.
So we made the idea more systematic—not as hidden metadata or text written for AIR, but as visible orientation for people.
We introduced a reusable page pattern called PageContext. On relevant product and question pages, it answers two simple questions:
- Who is this page for?
- What will it help them understand or do?
Mortgage, retirement, student-loan, Question Blueprint and AI Readiness pages began stating their audience and purpose directly.
AIR still scored Audience 0.
At that point, another wording change would have been the wrong response. The pages were clearer to people. The evidence existed in public, prerendered HTML. Rewriting it to satisfy a known assessment would have meant writing for the instrument rather than improving the website.
We needed to inspect the assessment itself.
The evidence was there
The investigation found something more important than a missed phrase.
AIR had discovered the relevant URLs through Wayli's navigation and sitemap. But its bounded acquisition process had not selected any of the pages containing the new PageContext evidence.
It had read:
- the homepage;
- About;
- the AI Readiness Example Report;
- Labs;
- Trust.
That sample was useful for Identity and Authority. It was much less representative of Wayli's products, questions and audiences.
The reason was architectural. Secondary-page selection had grown out of AIR's earlier Authority work, so it favoured organisation, trust and publication pages. Offer effectively relied on the homepage. Audience effectively relied on the homepage and an organisation page.
The problem was no longer simply:
AIR cannot recognise our audience wording.
It was:
AIR was not inspecting a representative enough sample of the website for the propositions it claimed to assess.
That distinction mattered. Loosening the Audience rules might have produced a more flattering result, but it would also have weakened the evidence standard we had just worked to establish.
We did not want AIR to infer more freely.
We wanted it to look in better places.
Found is not accepted
AIR V2 separates three events that can easily be confused:
Found != Accepted != Concluded
A page can be found without being read.
A page can be read without its contents being accepted as evidence.
Evidence can be accepted without being sufficient for a complete conclusion.
The new acquisition sequence became:
Discover broadly
↓
Classify likely page roles
↓
Choose representative coverage
↓
Acquire narrowly
↓
Evaluate the evidence using the existing deterministic rules
AIR now looks for a bounded mix of pages that may help investigate the organisation, its offer, its audience and its authority. A single page can cover more than one of those roles.
The page label is only a locator. A route called /customers, /services or /research receives no credit merely because of its name.
Page selection determines where AIR looks. Evidence rules determine what AIR can conclude.
Or, more simply:
Finding a page isn't the same as believing it.
Then we challenged the page limit
Improving selection raised another question.
Why should AIR read the homepage and four supporting pages? Why not three, seven or twenty?
The four-page limit had been inherited partly from the earlier Authority-oriented architecture. It was a practical bound, but not yet an empirically justified one.
So we tested acquisition depth across the maintained benchmark.
Fresh acquisition was attempted for 99 websites:
- 78 were assessable;
- 18 paused at preflight;
- 3 had technical acquisition failures.
For each assessable website, we measured whether each additional supporting page materially changed a proposition, confidence, uncertainty, score or recommendation. Another quote repeating an established point did not count as a material change.
The material changes by supporting-page position were:
- Page 1: 54 of 78 assessments, or 69.2%
- Page 2: 39 of 78, or 50.0%
- Page 3: 36 of 78, or 46.2%
- Page 4: 20 of 78, or 25.6%
- Page 5: 5 of 78, or 6.4%
- Page 6: 0
- Page 7: 0
The fifth supporting page was still useful. Pages selected beyond that boundary produced no material assessment change in this validation.
That supported the AIR V2 limit:
Homepage plus up to five representative supporting pages, stopping earlier when useful page roles are exhausted.
This does not mean websites need only six pages
A useful website may need tens, hundreds or thousands of pages.
AIR may discover many of them through navigation, sitemaps and public structure. Discovery does not mean reading every page.
The six-page maximum describes the representative sample AIR reads for this particular high-level assessment. AIR is investigating whether it can establish who an organisation is, what it offers, who it helps, why its information may be trusted and whether relevant public evidence is accessible.
It does not need to understand every product, location, guide or article to assess those propositions.
The benchmark supports a narrow claim: after AIR had selected representative page roles, pages beyond the boundary did not materially change the assessment in that sample.
Future evidence may justify revisiting the boundary. Bounded acquisition remains a methodology choice, not a natural law.
What we protected
A more representative sample is only useful if the conclusions remain defensible.
We replayed AIR's frozen Reference Website Panel controls after changing acquisition.
- Unsupported Audience assertions remained at zero.
- Negative Audience controls remained negative.
- Scoring weights did not change.
- Audience evidence rules were not loosened.
- No generative AI or probabilistic classifier was introduced.
The change was about where AIR looked, not what it was allowed to believe.
The new Wayli result
After AIR V2 was frozen, we assessed Wayli again from a fresh live acquisition.
- Identity: 100 in V1 → 100 in V2
- Offer: 100 → 100
- Audience: 0 → 60
- Authority: 100 → 100
- Information Access: 100 → 100
- Overall: 80 → 92
This is not simply a story about Wayli improving from 80 to 92.
The two scores belong to different methodology eras. Wayli's content evolved during the investigation, but V1 did not inspect many of the pages where that evidence lived. V2 used the same bounded philosophy to inspect a more representative sample of the website.
Audience is still only partially established, with Medium confidence.
AIR has not been tuned to give Wayli 100. Nor should it be.
The earlier 84 → 88 → 96 → 100 sequence also remains part of the record. Those were real results under the methods operating at the time. Preserving them makes it possible to see not only how Wayli changed, but how AIR's own standards changed too.
The current assessment and example report show the V2 result. The Wayli case study preserves the wider chronology, while the Methodology explains the current assessment at a higher level.
What we learned
We started by asking why Wayli scored zero for Audience.
The answer eventually led somewhere else.
The website had useful evidence. AIR had discovered the pages. But our own assessment was not giving those pages a fair opportunity to contribute.
Fixing that did not mean teaching AIR what to believe.
It meant improving where AIR looked, then leaving the evidence rules to decide what the pages actually supported.
That distinction now sits at the heart of AIR V2.