How We Score AI Visibility

An AI visibility score is worth exactly as much as the method behind it, so here is ours in full. Five areas, unevenly weighted: whether AI crawlers can reach the site, whether the content can be extracted, whether the structured data is valid, whether the business resolves to one unambiguous entity, and whether the site is ready for the way people now ask questions. What follows is the rubric itself, including what earns full marks in each area and what earns none.

Rated 5.0 from 20 Google reviews · A local SEO agency based in Haywards Heath

Why the rubric is public

Most scores in this field arrive as a number with a logo on it and no working. That is not a measurement, it is an assertion, and it is impossible to disagree with because there is nothing to disagree with. Publishing the bands means you can look at your own site, work out roughly where it sits before speaking to anybody, and tell us we have weighted something wrongly. Some clients have. Two of the bands below changed because of it.

The weighting, and why it is uneven

AreaWeightWhy that weight
AI crawler access25It is binary and it gates everything else. A site the engines cannot fetch scores nothing on the other four in practice, however good it is.
Content extractability25The difference between being quoted and being ignored, and the area most sites can move fastest.
Structured data25Not because it is an AI ranking factor, which Google says it is not, but because it is the cheapest way to remove ambiguity.
Entity clarity15Decisive when it is wrong, invisible when it is right, and largely off your own domain, which is why it is not weighted higher.
AI era readiness10The softest of the five and the most subjective, so it carries the least.

AI crawler access, out of 25

What is measured: whether the documented agents can actually retrieve a page, tested against the live site rather than read off robots.txt.

  • 25. Every documented search and retrieval agent gets a 200 and readable content on a plain request. Any blocks that exist are deliberate, consistent and aimed at training crawlers only.
  • 12. Reachable, but with friction: aggressive rate limiting, a consent layer between the request and the content, or key content that only appears after scripts run.
  • 0. A search or retrieval agent is blocked, by rule or by protection layer. This is usually accidental, and it is the single most common serious fault we find.

Content extractability, out of 25

What is measured: whether a bounded answer can be lifted from the pages that carry commercial weight, judged on those pages rather than on an average of the whole site.

  • 25. The answer is in the first sentence under a heading that matches the question, sections stand alone, and the scope and date are stated.
  • 12. The information is present and correct but has to be assembled from across the page, or the page opens with positioning before it answers anything.
  • 0. The commercial pages describe the service without ever answering a question somebody would ask.

Structured data, out of 25

What is measured: presence, validity and coherence of the core types, read from rendered HTML rather than from extracted text, which is where most automated checks go wrong.

  • 25. The core types are present, valid, non-conflicting, and the author and dates resolve to a real person and to real edits.
  • 12. Present but partial: a generic account as author, an automated dateModified that moves when a plugin updates, or overlapping types saying different things.
  • 0. None, or invalid enough that it will not be used.

Entity clarity, out of 15

What is measured: whether the business resolves to one thing everywhere it appears, on and off the domain.

  • 15. Name, address, phone number and description identical everywhere we can find them, and tied together in the markup.
  • 7. One or two stale records, usually an old address or a former trading name on a directory nobody has looked at in three years.
  • 0. Contradictory core details in more than one place a model can reach.

AI era readiness, out of 10

What is measured: whether the site behaves like a source rather than a brochure.

  • 10. Real named authors, honest dates, and at least one thing published that exists nowhere else.
  • 5. Competent, current, and entirely derivative.
  • 0. Undated, unattributed, and interchangeable with any competitor.

What the score cannot tell you

  • Whether an assistant currently names you. That requires asking the assistants and writing down what they say, on more than one day. It is the most useful check there is and it is not a score.
  • Whether your content is any good. Structure is measurable. Usefulness is not.
  • What anyone else says about you. Corroboration is most of the hard part and almost none of it is visible from your own domain.
  • Whether a low reading is real. Two of the competitor sites we assessed in August 2026 came back near the bottom on one attempt and near the top on another, with nothing changed in between; the low reading was throttling. Nothing is scored here on a single pass.

Working out roughly where you sit

Take the five areas above and mark yourself honestly, high, middle or low, in each. Most sites we see land between 45 and 65, and almost all of the gap is in the first two areas rather than the last three. Above 80 the technical foundations are sound and the remaining work is corroboration, which no rubric can measure. Below 50, start at the top of the list, because everything below it is moot if the engines cannot fetch the pages.

Frequently asked questions

Related guides

Sources

  1. Google Search Central. General structured data guidelines. Read 3 September 2026. The validity rules the structured data band is measured against.
  2. Google Search Central. Optimizing for Generative AI Features on Google Search. Updated 10 July 2026. Google’s statement that structured data is not required for generative AI search, which is why that area is weighted for clarity rather than as a ranking factor.
  3. OpenAI. OpenAI Crawlers and User Agents. Read 3 September 2026. The agent tokens the crawler access band is tested against.
  4. Perplexity. PerplexityBot and Perplexity-User documentation. Read 3 September 2026. Perplexity’s agents and published IP ranges, used in the access test.

Every figure on this page was read on the source page listed above. This page is reviewed monthly and the last checked date is updated when it is.

Frequently asked questions

Why is crawler access weighted as heavily as the content?

Because it is the only one of the five that can zero out the others. A beautifully written, perfectly marked up page that a retrieval agent cannot fetch contributes nothing at the moment an answer is written. It is also the fastest to fix, which makes it the best place to spend the first hour.

Yes, and this is the honest limit of any rubric. The score measures whether you are usable as a source. Whether you are chosen over somebody else depends on corroboration and on competition, neither of which is visible from your own domain. A high score removes the reasons not to cite you; it does not create a reason to.

Above 80 within a quarter is realistic for most sites, because the first two areas move quickly once someone is actually looking at them. Chasing 100 is not a good use of money. The last few points sit in the softest, most subjective area of the five.

Not because it matters less. Because most of it lives on other people’s websites, so it is the area a site owner can least directly control and the one where a score would be least fair as a judgement of the site itself. When it is wrong it is often the reason an assistant gets your details wrong, and we say so in the findings regardless of the weight.

CONTACT US

Ask us how your business shows up in AI search

A friendly 15 minute call, no obligation.

or call 01444 523105