Blocking Google-Extended removes you from AI answers while your rankings look untouched.
Bot and agent traffic passed half of all internet traffic, and most sites have never looked at which bots are hitting them or what their robots.txt is actually doing. The decisions here are real, they are reversible, and several of them are being made by accident in a file nobody has read since 2019.
Audit our crawler access-
Which bots reach you
- Named-agent log analysis. GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Google-Extended and the rest, separated.
-
What each one does
- Training, search indexing and on-demand fetch are three different jobs. Blocking the wrong one costs you visibility for no benefit.
-
The access policy
- A robots.txt that reflects a decision you made, written down, with the trade-off for each line stated.
-
What it costs you
- Bandwidth and origin load from crawlers that will never send you anything, separated from the ones that will.
The distinction that decides everything
Blocking training is not the same as blocking retrieval
These two decisions get made with one line in one file, and they have opposite consequences.
Blocking search crawlers while paying for AI visibility work is a combination we find more often than you would expect.
How a crawler access engagement runs
-
Log analysis
Weeks 1 to 2
- Traffic separated by named agent, with requests and bytes per bot
- What each crawler actually fetched, and how often it returned
- Crawlers hammering parameter or faceted URL space identified
- Origin cost attributed to bots that will never send you anything
-
Policy review
Week 2
- Current robots.txt read line by line, including rules that never apply
- Meta robots, X-Robots-Tag headers and CDN rules checked for conflicts
- Google-Extended and every AI token checked against what you intended
- Gap between what the file says and what you believe it says, written down
-
Decision and implementation
Weeks 2 to 4
- A policy you agree to, with the trade-off recorded per line
- Implemented in robots.txt, headers or at the CDN as appropriate
- Rate limits or blocks applied only where the log evidence supports them
- Change dated and documented so it is not a mystery in a year
-
Verification
Ongoing
- Logs re-read after the change to confirm the intended agents are behaving
- AI answer set re-run to confirm nothing dropped out
- New agents watched for as they appear, which they do regularly
- Reported in the portal alongside the rest of the program
Before you buy this
What crawler control can and cannot do
This is a plumbing service with real limits. Worth stating them before you spend anything.
What it does
- ✓Tell you which named AI agents actually reach your site, and what they cost you in origin load
- ✓Find the accidental blocks that removed you from AI answers without touching your rankings
- ✓Separate training, search indexing and on-demand fetch so each decision is made deliberately
- ✓Implement a policy you agree to and document why each line exists
- ✓Re-check the answer set afterwards so you can see the change rather than assume it
What it does not
- ✗Remove your content from a model that has already trained on it
- ✗Force an engine to cite you. Access decides eligibility, not selection
- ✗Guarantee every agent honours robots.txt. Most well-known ones do; not all traffic is well-known
- ✗Make llms.txt matter more than it does. Most llms.txt files are never fetched
- ✗Work properly without logs. Without them we are estimating, and we will say so
If your problem is that engines can read you and are not citing you, that is a GEO problem rather than an access one. See it for yourself →
How we report it
What the engines actually did, per engine
Crawler access decides whether an engine can read you. These two clients show what the resulting citation counts look like once it can, and how differently they land.
Delta Medical Labs
5,570 pages cited across AI assistants, July 2026
Eduverse
171 pages cited across AI assistants, August 2026
- 89.1%of the sites ChatGPT cites, Perplexity never touches for the same questionWellows, 804,058 answers, Sept 2025 to May 2026
- 79.6%of sources appear on one engine only22.7M citations across 1,146,483 questions, 2026
- 46xgap in brand citation rate between ChatGPT at 0.59% and Perplexity at 13.05%Study of 34,234 AI responses, 2026
Delta’s AI Overview count is 27 times its ChatGPT count. Eduverse’s top and bottom engines sit 18 pages apart. Same agency, same method, opposite shapes. Any single score we quoted you would have described neither.
Audit our crawler accessClient reviews
All reviewsWhat happens when
A crawler access engagement, week by week
Access changes show in logs within days and in answers within weeks. This is the honest sequence.
Log analysis
Traffic separated by named agent, with requests and bytes per bot, and what each one actually fetched.
Policy review
Your robots.txt read line by line, including rules that never apply because of ordering, plus header and CDN conflicts.
Decision and implementation
A policy you agree to, with the trade-off recorded per line, then implemented where it belongs rather than everywhere.
Verification
Logs re-read to confirm the intended agents behave, and the answer set re-run to confirm nothing dropped out.
New agents
They appear regularly. We watch for them rather than reviewing the file once a year.
Know what you are blocking
The named agents, and what each one does
Lumping these together is what produces the accidental blocks we find most often.
Questions about AI crawlers
Usually not all of them, and almost never the search and fetch ones. Blocking training is a defensible position for a publisher with a licensable archive. For most businesses it protects little and risks a lot, particularly if the block is written too broadly.
It removed you from generative grounding while leaving Googlebot and your rankings untouched. Nothing in Google’s tooling reports it, so it is usually found only when somebody checks the answers. It is reversible.
The major named agents document that they do, and in our logs they behave accordingly. Not all traffic identifies itself honestly. If a specific agent is a real cost, blocking at the CDN by behaviour is more reliable than asking politely.
As free hygiene, yes. As a paid line item, no. A study of 137,000 sites found the overwhelming majority are never fetched, and Google has said publicly you do not need one. We will set it up and we will not bill it as strategy.
They make it a measurement rather than an estimate. Without them we work from CDN reports or platform data and we tell you plainly that it is an approximation.
Crawler behaviour changes within days once the file is re-fetched. Whether you appear in answers again takes weeks, because the index has to be rebuilt with your pages in it.
