Journal 040

We Found a Score AIR Couldn't Award

Making AIR's scoring more transparent exposed an Audience score that could never be reached. Testing the problem led us to simplify the question rather than weaken the evidence standard.

5 min

Written by

Making AI Readiness scoring more transparent exposed a problem.

Audience could theoretically score 100.

In practice, AIR could never award it.

The current evidence engine could establish an intended audience and award 60, or fail to establish one and award 0.

There was no validated route to 100.

We could have made one.

Instead, we tried to find out what 100 should actually mean.

What would deserve 100?

The obvious answer was stronger evidence.

So we compared websites where independent review found the intended audience clearly established with websites where it was only partially established.

A distinction began to emerge.

Could AIR identify evidence that clearly connected an organisation or its main service to the people or organisations it was intended to help?

On a small controlled set, the answer looked promising.

The proposed rule reached 100% precision.

Then we tested it more widely.

It broke.

The statements were true

The problem wasn't bad evidence.

AIR found real statements about students, communities, clients, visitors and other groups.

But it could not always tell whether they described who the organisation was for or who one of its services, programmes or activities was for.

Rather than adding more words and exceptions, we tried to make that distinction structural.

We tested whether AIR could reliably identify what each audience statement was actually about: the organisation, a main service, an operational function, a programme or something else.

It worked in many cases, but not enough.

Before the final test, we set a threshold: the method needed to retain at least 15 of 21 independently defensible examples.

It retained 12.

So we stopped.

We could probably have added more rules.

That wasn't the test.

We asked a simpler question

The investigation had shown that AIR was much better at answering something narrower:

Can AIR establish an intended audience from the public evidence it acquired?

That was already close to what the evidence engine was actually doing.

So we tested that question instead.

Across the independently reviewed Reference Panel:

  • Precision: 100%
  • Recall: 91.3%
  • Specificity: 100%
  • False audience establishments: 0

That became the Audience question for the new methodology.

What 100 means now

Audience is now binary.

If AIR establishes a qualifying intended audience from the evidence it acquired:

100

If it cannot:

0

That does not mean AIR is 100% confident.

Score and confidence answer different questions.

Score asks: did the evidence satisfy the defined assessment test?

Confidence asks: how strongly does the acquired evidence support AIR's understanding?

So an assessment can legitimately say:

Audience established — 100/100 Moderate confidence

It means AIR established an intended audience.

It does not mean every possible audience was identified, the website is perfect or AIR is 100% certain.

The methodology changed, not the past

This changes the scoring contract, so it belongs to a new methodology version.

Earlier assessments remain what they were.

Wayli's previous 92/100 result remains a valid result under AIR V2.

A new assessment may produce a different score under the newer methodology even if the website itself has not changed.

That is why the methodology version belongs with the result.

The current assessment and Evidence Report show the V3 result. The Methodology defines the current rules, while the Wayli case study preserves the earlier results alongside it.

What we kept

The failed experiments weren't wasted.

They exposed useful future questions around:

  • primary proposition scope;
  • conflicting evidence;
  • stronger Audience understanding;
  • decision provenance.

We also built more of the infrastructure needed to record how AIR reaches its decisions.

Those ideas can return in a future methodology version if the evidence supports them.

They don't belong in the current score simply because we would like a more sophisticated model.

What we learned

We started by trying to define what deserved 100.

The benchmark eventually persuaded us that we were asking AIR to make a distinction it could not yet make reliably enough.

So we changed the question.

Not to make the score easier.

To make the claim narrower.

A benchmark shouldn't exist to prove your rule works. It should be capable of persuading you to abandon it.

Sometimes improving an assessment means asking a better question.