Journal 019

We Asked AI to Test 100 Businesses

Testing one website taught us about Wayli. Testing one hundred began to teach us about AI.

5 min

Written by

We Asked AI to Test 100 Businesses

Testing Wayli had taught us something important.

Improving the clarity of a website appeared to improve AI's understanding of it.

But there was an obvious problem.

One website proves almost nothing.

One business.

One report.

One result.

That's not evidence.

It's the beginning of an investigation.


A Bigger Question

So we asked a bigger question.

What happens when the same assessment is applied to one hundred different organisations?

Not one carefully chosen example.

One hundred genuine public websites.

Different industries.

Different sizes.

Different ways of explaining what they do.


We Weren't Looking For Winners

This wasn't about creating another league table.

In truth, we weren't especially interested in who scored highest.

We were looking for patterns.

  • Were businesses consistently strong in some areas?
  • Were the same gaps appearing repeatedly?
  • Could AI confidently identify who organisations served?
  • Could it explain why those organisations should be trusted?

Individual scores mattered.

Recurring patterns mattered much more.


We Learned About Ourselves Too

Running one hundred assessments also tested the Wayli Understanding Engine.

Most websites completed successfully.

Some didn't.

Surprisingly...

...those failures turned out to be useful.

A failed assessment isn't necessarily a failed experiment.

Sometimes a website blocks automated access.

Sometimes it depends on JavaScript rendering.

Sometimes it exposes a weakness in our own acquisition process.

Every unexpected result became another opportunity to improve the engine.


The First Pattern

Although one hundred organisations is still a relatively small benchmark, one observation stood out.

Many businesses appeared less AI-ready than we expected.

Not because they were poor businesses.

Not because their products lacked quality.

But because their websites often assumed knowledge that AI couldn't safely infer.

The clearer the evidence, the stronger the understanding.

That idea kept appearing.


What One Hundred Really Taught Us

Testing one hundred websites won't tell us everything.

The internet is far too diverse for that.

But it gives us something much more valuable than opinion.

It gives us evidence.

And evidence allows better questions to be asked.

The benchmark wasn't valuable because it produced one hundred scores.

It was valuable because it began to reveal repeatable patterns.

Those patterns would go on to shape every version of the Wayli Understanding Engine that followed.