Can all registered uses please login, even just for a few minutes..
It helps build a picture where our "good traffic" is coming from..
Thanks :)

AI crawlers: where do you all stand on this?

This chat forum will appear to guests and bots. IE will appear on search engines etc. If you do not want your appearing publicly then please use the original chat forum.
Post Reply
User avatar
exxos
Site Admin
Site Admin
Posts: 28972
Joined: Wed Aug 16, 2017 11:19 pm
Location: UK
Contact:

AI crawlers: where do you all stand on this?

Post by exxos »

Something I keep going back and forth on, and I am curious what everyone else thinks, because I do not think there is a clean answer.

Over the past few days I have been digging through the server logs and the picture is fairly stark. One crawler alone, KeenableBot (keenable.ai, a search index for AI agents that came out of stealth in August with a $26 million funding round), was doing over half a million requests a day, more than 60% of everything hitting the box, from a couple of dozen Google Cloud addresses that rotate daily. It had read robots.txt exactly once and then ignored the one instruction in it, running at ten times the rate it was asked to, with individual addresses recorded doing thirty requests in a single second. It also identified itself as SleepBot until the 11th, then changed its name and carried on doing exactly the same thing. That one is now blocked, and I would rather not have had to. Meta's Facebook crawler has been blocked here for a while for much the same reason.

The other thing that gets me is the sheer pointlessness of a lot of it. This is not a huge forum. Anything crawling it sensibly could have the whole lot inside an hour and be done. Instead some of them just loop, round and round the same forums, the same topics, day after day, requesting content that has not changed since the last time they asked for it an hour ago. There are perfectly good mechanisms built into the web for saying "this has not changed since you last looked", and they were designed for exactly this, but plenty of these crawlers do not appear to use any of them. So the same bytes get sent thousands of times over, which is a waste of my bandwidth and, oddly enough, a waste of theirs too. Nobody wins.

And yet. I actually want AI crawlers to have this content. Half the point of this place is that somebody's write-up of a dodgy floppy drive in 2009 is still useful to somebody in 2026. If an AI has read it and can hand that answer to the next person with the same fault, that is the content doing exactly what it was written to do, just by a different route. I use Claude and Grok myself, daily, for Atari work and for the forum's own software, and I have had them hand me back things that almost certainly came off this very forum. So in a roundabout way the content does come back around and help someone, including me. I pay for those services and I have no problem with that, given the compute involved is not exactly free either. That is at least a recognisable arrangement: they built something useful and I am paying to use it.

Which is also where the sting is. I am paying them, and they took the raw material from here for nothing.

So the crawling itself is not really my objection. It is two other things.

The first is the sheer rudeness of some of it. There is a difference between reading a site and hammering it. A crawler that reads robots.txt, respects a one second delay and quietly works through the sitemap costs me nothing worth mentioning. A crawler opening connections as fast as it physically can, ignoring every signal it was given, is just taking the mick. I end up blocking those not because I object to what they want but because of how they go about getting it.

There is a wrinkle in robots.txt worth spelling out too, because it is not as useful a tool as people assume. The crawl delay is per connection, not per company. So a crawler can honour it to the letter, wait its full second between requests, and still hammer you into the ground simply by running a thousand of itself at once from a thousand different addresses. Every individual one is behaving impeccably. The site still sees a request every millisecond. I have watched exactly this happen here, with addresses that rotate daily and sometimes shift data centre entirely, so there is nothing stable to point a rule at. At that point robots.txt is not really doing the job it looks like it is doing, and you are down to two options: absorb it, or block it.

The second is harder to shrug off. Everything on here was written by people, for free, to help other people. It gets hoovered up wholesale, and then sold, sometimes by companies who charge their own customers for access to a search index built out of it. Nobody who wrote any of it sees a penny, nobody asked first, and the only person whose costs go up is whoever is paying the hosting bill. That is not AI being evil, it is just a plainly one-sided arrangement, and it would be a lot easier to stomach if any of them ever said thank you, or throttled themselves, or offered anything at all in return.

The EU has been drafting rules about opt-outs and training data disclosure, which is well intentioned, but it feels rather like arriving with a fire extinguisher once the house has burnt down. The entire internet has already been indexed several times over. Whatever rules arrive now apply to the next round, not to anything already taken.

So I am genuinely torn. Blocking crawlers protects the server and makes a point, but it also means this forum quietly vanishes from the tools an increasing number of people use to find answers, which does nobody here any favours. Letting them all in means paying to feed other people's products.

Where do you lot stand? Does it bother you that things you have written here are almost certainly sitting in a few different models by now, or is that just the price of putting anything on the internet in the first place? And does it change your view at all if the thing that swallowed it is at least handing useful answers back out the other side?
Post Reply

Return to “CHAT FORUM PUBLIC”