Who you let in
Nine crawlers, asked for one at a time.
GPTBot · ClaudeBot · PerplexityBot · Bingbot · Applebot · Amazonbot
Google-Extended · CCBot · Bytespider
sites block an AI crawler
and almost none of them meant to
1 in 3
Before ChatGPT can quote you, something has to fetch your page and get words back. Plenty of sites hand it an empty box instead, and never find out.
6,543 readable of 356,603 characters served
worked example · not a live scan
Every answer is a plain yes or no you can check yourself. Nothing is bundled into a grade, because a grade hides which one is hurting you.
Nine crawlers, asked for one at a time.
GPTBot · ClaudeBot · PerplexityBot · Bingbot · Applebot · Amazonbot
Google-Extended · CCBot · Bytespider
sites block an AI crawler
and almost none of them meant to
1 in 3
We fetch the page the way a crawler does, with no browser. Then we count what is still there.
of pages come back empty
they only fill in once JavaScript runs
36%
Your HTML against your actual words. Navigation, fonts and scaffolding are not what gets quoted.
chrome on one docs page
392KB of HTML, almost none of it readable
97%
One H1, headings in order, a title that matches, alt text that says something. The boring half that decides whether a sentence can be lifted.
is not the same as readable
one site returned 17,224 bytes of headers and no page
200
Neither of them was lying. One stopped reading early, the other read to the end, and both called the answer the same word. That is why this page hands you counts you can check instead of a number out of a hundred.
Cut the HTML off at a fixed size, then counted what was left.
Read the whole document, then counted.
It is the first thing anyone is told to add, so it is worth being straight about. Three separate studies looked at it. None of them are ours, and all three landed in the same place.
We report whether yours exists. We never score you on it, and we will not sell you one.
Six answers, including the two that talk you out of trusting a number.
Yes, and there is no account. Type a URL, get the page back. We do not store the URL, we do not email you, and there is nothing to upgrade to at the end of it.
It asks for your page nine times, once as each named crawler, with no browser and no JavaScript. Then it counts what came back and compares that against what a person would have seen. Everything it reports is something you can verify yourself with curl.
No, and anyone telling you otherwise is selling something. This checks whether you are reachable and readable. Being quoted is a separate question that depends on what you actually say and who else says it. Reachable is the floor, not the ceiling.
Because it measured something else. We ran the same page through two of them on the same afternoon and got 1.8% and 5%. One truncated the HTML before counting, the other read all of it. Neither was lying. This is why we show you the raw counts instead of a score out of 100.
Most sites are built so the browser assembles the page after it arrives. A crawler that does not run JavaScript gets whatever was in the file, which on a lot of modern sites is an empty div and a script tag. That gap is the single most common thing this finds.
It says so, and it is left out of the result rather than scored zero. A report that punishes a perfectly good site for not having something it never needed is worse than no report.