AI crawlers are inflating your cloud bill. Here's the fix.
Your traffic graph went up. Nobody visited. Both of those things are true at the same time, and the second one is exactly why the first one costs you money.
You’re already in the minority on your own site
More than 53% of all web traffic was automated in 2025 — up from 51% the year before — according to Imperva’s 2026 Bad Bot Report. Human activity is down to 47% and still falling. That’s not a prediction about the future of the web. That’s the web you’re serving right now, this afternoon.
And here’s the part that catches almost everyone: your analytics never told you. Google Analytics, Matomo, any JavaScript tag — they only count visitors who run your JavaScript, and bots overwhelmingly don’t. So your dashboard drew a calm, flat, human-shaped line while your server was quietly serving a completely different and much larger website. The two systems weren’t disagreeing. They were measuring different things, and only one of them was connected to your bill.
The access log is the only honest record you have of what your infrastructure actually did.
Your analytics measure your audience. Your bill measures your traffic. Nobody told you those stopped being the same number.
The honest math (it’s not the bandwidth)
Most articles on this topic want you to panic about bandwidth. Let’s do the arithmetic instead, because the real answer is more useful than the scary one.
Say two million bot requests a month at 80 KB a page. That’s about 160 GB of egress — roughly $14 at AWS’s standard rate of about nine cents a gigabyte. Nobody’s business dies from $14. If you run a small static site, bot traffic is a rounding error and you can stop reading.
The bill shows up when your pages aren’t static, and it shows up on four meters nobody thinks to check:
- Compute and database. Every crawler request that reaches a real page render is a set of database queries and some CPU. A human browses maybe six pages. A crawler methodically walks everything — the archive nobody reads, every tag page, every filter combination, every paginated listing. That is precisely the traffic your cache was never warmed for, so a disproportionate share of it hits your origin cold.
- Autoscaling. If your infrastructure scales on load, it scales on their load too. You provisioned for a crowd that isn’t buying anything.
- Per-request fees. CDN and WAF pricing is charged per request. The bot’s request costs exactly what a customer’s request costs, and it converts at zero.
- Logs. Every one of those requests writes a log line you pay to ingest, pay to store, and pay again to scan later. Bot traffic doesn’t just cost you at the door — it silently inflates the meter you’ll read next quarter.
None of those four are dramatic on their own. That’s what makes this the same species as the $100k already hiding in your bill: it’s not a spike, it’s a tilt. Nothing looks broken. The number is just quietly bigger than it should be, in a category nobody owns.
The ratio that should decide it for you
Here’s the fact that reframes this from a technical annoyance into a business decision.
Cloudflare started publishing crawl-to-refer ratios — how many pages a platform’s crawler takes for every visitor it sends back. For the week of June 19–26, 2025, Anthropic’s crawler made roughly 70,900 HTML page requests for every single referral it sent to a site. Mistral, at the other end, sent about ten referrals for every page it crawled.
Sit with the first number for a second. The old bargain with search engines was legible: Google scanned your pages a couple of times and sent you a visitor. You paid for the crawl and got traffic — that’s rent, and it’s fair. Cloudflare’s own data shows that by mid-2025, training drove nearly 80% of AI crawling, and GPTBot’s share of requests grew from 4.7% to 11.7% in a year.
A 70,900:1 ratio isn’t a bargain. It’s a one-way transfer, and you’re paying the freight.
Google took two pages and sent you a customer. That was rent. This is not rent.
To be fair about it: those crawlers may still be worth serving, if being cited inside an AI answer is where your buyers now are. That’s a real strategy and I’m not going to talk you out of it. But it should be a decision you made with the ratio in front of you, not a default you drifted into because nobody looked at the log.
What to actually do, in order
Read the log first. Fifteen minutes, no purchases. Pull last month’s access log and group the requests
by user-agent. You’re looking for GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider, CCBot,
PerplexityBot, Amazonbot — and, just as importantly, for the large pile of requests wearing a normal
browser’s user-agent that came from one hosting provider’s address range. Now you know your split. Almost
everyone is surprised, and the surprise is usually the point where this stops being theoretical.
Cache before you block. This is the highest-leverage move and the one people skip because blocking feels more satisfying. A crawl that hits a warm cache costs you a file read; the same crawl hitting your origin costs you a page render, a database round-trip and a log line. Caching cuts the bill without you having to predict which bots deserve to live — and it makes your site faster for the humans in the same move.
Then decide per-bot, using the ratio. Not “block all bots” — Googlebot still pays rent, and so might some AI crawlers. Block the ones that take and never give.
Use robots.txt, but know what it is. It’s a request, not a fence. The well-behaved crawlers honor it
— which is genuinely most of the big named ones — and the ones you’re most worried about ignore it entirely.
Set it correctly anyway, because it’s free and it makes the polite majority go away.
Enforce at the edge for the rest. Block or challenge by user-agent and network at your CDN or WAF, so the request dies before it reaches anything you pay for per-invocation. A block that happens at your origin already cost you most of the money.
And know that a price tag now exists. Cloudflare’s pay per crawl lets a site answer a crawler with
402 Payment Required and a price instead of a free page. Whatever you think of the economics, it’s a
useful signal: the industry has now formally admitted that serving a crawler is a cost, not a courtesy.
The same mistake, one layer up
If this feels familiar, it should. It’s the same shape as your support bot being someone’s free ChatGPT, and the same shape as an agent with a goal and no budget: something automated found a resource you left open, and your bill is the only place it shows up. The pattern underneath all three is the unlimited-liability default — you agreed to serve whatever asks, and you agreed to be billed for all of it, and nobody made you sign anything.
The web quietly became majority-machine. Your infrastructure noticed. Your dashboard didn’t.
If your traffic looks healthy but your bill looks wrong, the answer is usually in the access log, and it usually isn’t human. Send me a slice of your log and your last bill and I’ll show you what share of what you’re paying for is machines, and which of them are sending anything back. Every message comes straight to me — I read and reply to each one myself, usually within a business day, and what readers send shapes what I build next. Send it over — free, within a day.