Can all registered uses please login, even just for a few minutes..
It helps build a picture where our "good traffic" is coming from..
Thanks :)

AI crawlers: where do you all stand on this?

This chat forum will appear to guests and bots. IE will appear on search engines etc. If you do not want your appearing publicly then please use the original chat forum.
Post Reply
User avatar
exxos
Site Admin
Site Admin
Posts: 28978
Joined: Wed Aug 16, 2017 11:19 pm
Location: UK
Contact:

AI crawlers: where do you all stand on this?

Post by exxos »

Something I keep going back and forth on, and I am curious what everyone else thinks, because I do not think there is a clean answer.

Over the past few days I have been digging through the server logs and the picture is fairly stark. One crawler alone, KeenableBot (keenable.ai, a search index for AI agents that came out of stealth in August with a $26 million funding round), was doing over half a million requests a day, more than 60% of everything hitting the box, from a couple of dozen Google Cloud addresses that rotate daily. It had read robots.txt exactly once and then ignored the one instruction in it, running at ten times the rate it was asked to, with individual addresses recorded doing thirty requests in a single second. It also identified itself as SleepBot until the 11th, then changed its name and carried on doing exactly the same thing. That one is now blocked, and I would rather not have had to. Meta's Facebook crawler has been blocked here for a while for much the same reason.

The other thing that gets me is the sheer pointlessness of a lot of it. This is not a huge forum. Anything crawling it sensibly could have the whole lot inside an hour and be done. Instead some of them just loop, round and round the same forums, the same topics, day after day, requesting content that has not changed since the last time they asked for it an hour ago. There are perfectly good mechanisms built into the web for saying "this has not changed since you last looked", and they were designed for exactly this, but plenty of these crawlers do not appear to use any of them. So the same bytes get sent thousands of times over, which is a waste of my bandwidth and, oddly enough, a waste of theirs too. Nobody wins.

And yet. I actually want AI crawlers to have this content. Half the point of this place is that somebody's write-up of a dodgy floppy drive in 2009 is still useful to somebody in 2026. If an AI has read it and can hand that answer to the next person with the same fault, that is the content doing exactly what it was written to do, just by a different route. I use Claude and Grok myself, daily, for Atari work and for the forum's own software, and I have had them hand me back things that almost certainly came off this very forum. So in a roundabout way the content does come back around and help someone, including me. I pay for those services and I have no problem with that, given the compute involved is not exactly free either. That is at least a recognisable arrangement: they built something useful and I am paying to use it.

Which is also where the sting is. I am paying them, and they took the raw material from here for nothing.

So the crawling itself is not really my objection. It is two other things.

The first is the sheer rudeness of some of it. There is a difference between reading a site and hammering it. A crawler that reads robots.txt, respects a one second delay and quietly works through the sitemap costs me nothing worth mentioning. A crawler opening connections as fast as it physically can, ignoring every signal it was given, is just taking the mick. I end up blocking those not because I object to what they want but because of how they go about getting it.

There is a wrinkle in robots.txt worth spelling out too, because it is not as useful a tool as people assume. The crawl delay is per connection, not per company. So a crawler can honour it to the letter, wait its full second between requests, and still hammer you into the ground simply by running a thousand of itself at once from a thousand different addresses. Every individual one is behaving impeccably. The site still sees a request every millisecond. I have watched exactly this happen here, with addresses that rotate daily and sometimes shift data centre entirely, so there is nothing stable to point a rule at. At that point robots.txt is not really doing the job it looks like it is doing, and you are down to two options: absorb it, or block it.

The second is harder to shrug off. Everything on here was written by people, for free, to help other people. It gets hoovered up wholesale, and then sold, sometimes by companies who charge their own customers for access to a search index built out of it. Nobody who wrote any of it sees a penny, nobody asked first, and the only person whose costs go up is whoever is paying the hosting bill. That is not AI being evil, it is just a plainly one-sided arrangement, and it would be a lot easier to stomach if any of them ever said thank you, or throttled themselves, or offered anything at all in return.

The EU has been drafting rules about opt-outs and training data disclosure, which is well intentioned, but it feels rather like arriving with a fire extinguisher once the house has burnt down. The entire internet has already been indexed several times over. Whatever rules arrive now apply to the next round, not to anything already taken.

So I am genuinely torn. Blocking crawlers protects the server and makes a point, but it also means this forum quietly vanishes from the tools an increasing number of people use to find answers, which does nobody here any favours. Letting them all in means paying to feed other people's products.

Where do you lot stand? Does it bother you that things you have written here are almost certainly sitting in a few different models by now, or is that just the price of putting anything on the internet in the first place? And does it change your view at all if the thing that swallowed it is at least handing useful answers back out the other side?
User avatar
stephen_usher
Site sponsor
Site sponsor
Posts: 7535
Joined: Mon Nov 13, 2017 7:19 pm
Location: Oxford, UK.
Contact:

Re: AI crawlers: where do you all stand on this?

Post by stephen_usher »

*IF* the crawlers were not anti-social, slowly trawled sites and respected robots.txt like other crawlers I would have no issue with them. They're just another indexing engine.

*HOWEVER* the way they have been written makes them an Internet denial of service attack and hence I would consider them in the same boat as other DDoS malware on the 'Net. i.e. criminal attack bots.
Intro retro computers since before they were retro...
ZX81->Spectrum->Memotech MTX->Sinclair QL->520STM->BBC Micro->TT030->PCs & Sun Workstations.
Added code to the MiNT kernel (still there the last time I checked) + put together MiNTOS.
Collection now with added Macs, Amigas, Suns and Acorns.
User avatar
PhilC
Moderator
Moderator
Posts: 7582
Joined: Fri Mar 23, 2018 8:22 pm

Re: AI crawlers: where do you all stand on this?

Post by PhilC »

I think I agree with Stephen. If their action causes the needless consumption of resources by not following instructions, then they should be blocked.
If it ain't broke, test it to Destruction.
User avatar
exxos
Site Admin
Site Admin
Posts: 28978
Joined: Wed Aug 16, 2017 11:19 pm
Location: UK
Contact:

Re: AI crawlers: where do you all stand on this?

Post by exxos »

Yeah, look at the traffic before and after blocking just that 1 bot..

Capture.PNG
Capture.PNG (117.09 KiB) Viewed 29 times
User avatar
mrbombermillzy
Moderator
Moderator
Posts: 2414
Joined: Sun Jun 03, 2018 7:37 pm

Re: AI crawlers: where do you all stand on this?

Post by mrbombermillzy »

There was a recent news article about researchers trying to solve a long standing mathmatics/fluid dynamics problem, with AI beating them to the solve.

Apparently, it used their online research data to do it: https://finance.biggo.com/news/b05bc0db ... 3da78adc03

So bearing this in mind, if this trend continues then people are going to become more and more reluctant to post any of their work online.

Its one thing being happy to contribute to open source projects - even with completely relaxing the relevant licences involved, but quite another to have your work unceremoniously ripped off and used as a means to charge customers for a 'higher tier' of information/data.

Back to your question though:

I believe keeping the data trawlers at bay will actually enhance/improve the forum membership, as people will - if interested in the relevant subjects detailed here - join up to find the answers and it will hopefully stop users from withholding the fruits of their blood, sweat and tears who are unhappy with the above situation.

Perhaps eventually the 'middle road' here can be achieved, if at some future point 'ethical bot etiquette' guidelines are drafted and adhered to by the companies involved. We can only hope.

Until then, I say pull the plug. :roll:
User avatar
exxos
Site Admin
Site Admin
Posts: 28978
Joined: Wed Aug 16, 2017 11:19 pm
Location: UK
Contact:

Re: AI crawlers: where do you all stand on this?

Post by exxos »

One related thing which just popped in my head, if people stop publishing stuff on the Internet, then what trains the AI models in the future ?

All the AI models depend on worldwide amount of data, and if that data stops, because everyone has shifted to using AI to solve problems, and nobody writes articles any more, at what point does it all backfire.... Does AI models them become "stale" and then unable to solve "new" problems as it doesn't have worldwide user base of data to draw from anymore..
User avatar
mrbombermillzy
Moderator
Moderator
Posts: 2414
Joined: Sun Jun 03, 2018 7:37 pm

Re: AI crawlers: where do you all stand on this?

Post by mrbombermillzy »

exxos wrote: Mon Sep 14, 2026 4:43 pm Does AI models them become "stale" and then unable to solve "new" problems as it doesn't have worldwide user base of data to draw from anymore..
It must surely reach a saturation point at some stage - with or without people providing data, so I guess 'yes' is the answer to that. :)
Post Reply

Return to “CHAT FORUM PUBLIC”