A new wave of crawlers has been hitting websites at a scale most server logs weren’t built to easily separate from regular traffic, and a lot of site owners genuinely have no idea it’s happening or what’s being done with what gets pulled.
Who’s actually crawling
Beyond the familiar search engine bots, there’s now a growing list of AI-company crawlers indexing content to train models or to answer queries directly inside a chat interface, often bypassing the click to your site entirely. Some respect robots.txt. Some historically haven’t, until public pressure made them start.
The traffic you can’t see
A visitor asking an AI assistant a question your article answers may get a summarized version of your content with no visit, no ad impression, and no line in your analytics. The value transfer is real. The attribution mostly isn’t, at least not yet.
What site owners are actually doing about it
Some are blocking these crawlers outright in robots.txt, betting the training-data value isn’t worth the lost visibility. Others are doing the opposite, optimizing content specifically to be cited by AI answers, on the theory that being the quoted source beats being invisible entirely. Both are reasonable bets. Neither has a long enough track record yet to call a winner.
What’s not reasonable is the current default for most small sites: not checking server logs at all, and finding out a year later that a third of “organic” traffic patterns changed and nobody knows why.
Check your logs for crawler user-agents you don’t recognize. It takes twenty minutes and it’s the only way to know who’s actually reading your site before your visitors do.
Comments are closed on this one.