Rendered at 17:07:23 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
simonw 16 hours ago [-]
The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini:
> Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).
dannyw 13 hours ago [-]
Good. Google's approach here is manifestly predator, unfair, and IMO illegal. They deserve to be in court for this behaviour, and mandating owners give consent for AI training or drop out of Google; which is just a non-starter because they're a search monopoly.
That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a trace for damages.
AnthonyMouse 7 hours ago [-]
The irony here is that the people blocking all other crawlers are the ones shoring up their monopoly. If you can't block Googlebot because you need the search traffic but you block everybody else so that nobody other than Google can index your site, how do you expect to ever get any search traffic that isn't from Google?
tempest_ 2 hours ago [-]
Search traffic is nose diving due to LLM use. The fundamental calculus with google is you install GA and it helps your SEO has changed and google is riding out what will eventually wither as people start to reevaluate the trade off.
troyvit 12 hours ago [-]
I think it's bad, because everybody is desperate to hold onto every last bit of google search traffic they can, so they're going to allow training to do so. Google's predatory, unfair and illegal actions will continue as they have with a few $100 million slaps in the wrist from the EU and a few more white house dinners for their CEO.
Deathmax 3 hours ago [-]
Except that officially, that is not what they do? It's a dick move not to put the two usecases under separate user agents, but their documentation says you're free to block Google-Extended via robots.txt which is used for training and grounding, while still being included in the search index.
Exclusion from grounding does mean that your site won't get sourced in the AI overview, but I'm not sure what the click through rates are like on those.
mysterydip 3 hours ago [-]
AI crawlers, famous for respecting robots.txt ;)
Scroll_Swe 5 hours ago [-]
EU will write some strongly worded letter.
Saying this as a European who is pro EU.
Why should they?
PunchyHamster 10 hours ago [-]
I don't think it's the Google bot DDOSing people's infrastructure for AI training...
9 hours ago [-]
xhi 9 hours ago [-]
[dead]
Razengan 5 hours ago [-]
This had me in disbelief since the minute I saw it: Google's "AI overview" presumably trained on content from other websites, disincentivizes users from clicking through to those websites..
How is that not conflict of interest??
paulddraper 2 hours ago [-]
> mandating owners give consent for AI training or drop out of Google
Huh?
Search and AI are hand-in-hand.
They both rely on embeddings. (Unless you still do keyword-only search, but that's not as good.)
inigyou 6 hours ago [-]
How come you say "Good." and you are near the top of the comments but I say "Good." and get flagdead?
dijksterhuis 5 hours ago [-]
it’s likely to do with the fact that the parent comment here laid out a thoughtful basis / argument for their “good”, providing some detail to their justification for it.
that’s just my take/feedback, take it or leave it. i won’t be engaging further as i already feel i’m going against the site guidelines with this!
jofzar 16 hours ago [-]
We had googlebot blast a random customer system and almost cause an outage, this is when I first learnt that google will use it for AI training also. It's honestly kind of frustrating also because you then search on it and theres (was) nothing on how you are meant to "correctly" tell google to fuck off, and not use it like that.
If you're using Cloudflare, set up a security rule to block requests that have "Googlebot" in the UA and are not recognised by CF as a real bot.
20k 15 hours ago [-]
Google's web scraping functionality has been acting as a ddos for more than two decades. I've seen literally hundreds of reports of them attacking websites and taking them down, where there's nothing you can do but accept the traffic, or get delisted
This is unfortunately nothing new. There's no correct way to tell them to fuck off, they do not care, and they never will do. People have even taken them to court over this
weird-eye-issue 12 hours ago [-]
If a site cannot handle traffic from the real Googlebot that is a serious issue with the site itself since it's actually pretty conservative
Also I should note there are lots of fake Googlebots...
remus 11 hours ago [-]
Indeed, I've got a site which gets a lot of bot traffic and google bot is pretty sensible compared to a lot of other mainstream bots.
weird-eye-issue 8 hours ago [-]
Yes, it really does not make that many requests. In fact lots of site owners struggle with having it not crawl and index their site enough
20k 5 hours ago [-]
It is mostly, but it doesn't take a lot of googling to find sites getting ridiculous amounts of traffic from googlebot on google IPs. Its one of the most common complaints about google's search indexing
weird-eye-issue 4 hours ago [-]
Lots of people abuse Google Cloud to get a "Google IP" for a fake Googlebot. Why don't you show me a single screenshot from Google Search Console showing a high number of requests to a site that would be counted as a DoS? All requests from the official Googlebot are logged there so if it's such a common problem it must be very easy for you to show me this.
motbus3 7 hours ago [-]
Don't that feel like a threat to businesses who dare to avoid their content being stolen?
miohtama 12 hours ago [-]
People will use something for search and something needs to index pages, either for LLM or old school search engine.
inigyou 16 hours ago [-]
[flagged]
Cider9986 15 hours ago [-]
Why should I use something other than Cloudflare pages for a simple app landing page?
inigyou 8 hours ago [-]
Because you value the internet being decentralised.
ipaddr 14 hours ago [-]
Because your viewer/customer base will be reduced.
Cider9986 10 hours ago [-]
It would have a domain. It affects it even then?
inigyou 8 hours ago [-]
Well, now Google won't be able to see it.
ajmurmann 15 hours ago [-]
Why is this?
ceejayoz 15 hours ago [-]
It's a planet-scale MITM?
TurdF3rguson 14 hours ago [-]
It's a cache. My tiny websites couldn't survive getting hammered by AI bots without them.
inigyou 8 hours ago [-]
Are you sure? Have you tried, or did Cloudflare just tell you that?
dbbk 15 hours ago [-]
So you're against all CDNs?
14 hours ago [-]
fc417fc802 14 hours ago [-]
A CDN doesn't necessarily have to perform a MitM. We really need more nuanced terminology to distinguish the various approaches.
gruez 14 hours ago [-]
Right, but practically speaking all CDNs are MITMs. If you're against cloudflare you should be against cloudfront, akamai, etc. as well.
inigyou 8 hours ago [-]
Cloudflare is egregiously bad because of its marketing strategy. It tried to get everyone with any small website to use it, by selling a vague notion of security and charging no monetary price, and it worked. They'll even sell you a domain name to increase lockin. Many people recommend getting domains from cloudflare because apparently they're cheap.
Akamai, Fastly, etc only take big customers who know what they're doing. You need to sign a proper contract with them. They aren't low-friction.
sandeepkd 14 hours ago [-]
Ideally yes, the TLS termination does not need to happen for caching purposes. Challenge is that in practice every business wants to be sticky and try to provide more functionalities which do require TLS termination. Most people either trust CDN's or they do not understand MitM so it does not concerns them. Plus they are getting certificate management and DDOS prevention capabilities.
edaemon 14 hours ago [-]
How would they cache and serve responses without decrypting the traffic?
tekacs 16 hours ago [-]
> For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.
It's kind of exhausting seeing Cloudflare playing both sides of the arms race.
I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this.
> This also lines up the incentive model we want to foster. Losing trusted status across the more than 20% of web domains that sit behind Cloudflare is a deterrent with teeth. Trust becomes something you can carry with you, and something you can lose.
And even more so, LLM language aside, fun and fascinating to see them flagrantly calling out their position here as if it's a positive.
usef- 14 hours ago [-]
How are they playing both sides? I thought their scraping products were also about having it behave and not take down systems
inigyou 8 hours ago [-]
Their scraping products are mostly about paying them money to not be blocked.
siva7 7 hours ago [-]
So Google has to pay Cloudflare 10B$ to get Googlebot moved to their default Allowlist otherwise Google would loose access to a significant portion of the internet? Genius move.
inigyou 6 hours ago [-]
Yes literally this. Every large corporation gatekeeper does this sort of thing, it's why we have to militantly push for decentralisation by avoiding these corporations. It's not morally different from Google forcing your site to be AI-scrapeable or be kicked off Google Search. It's why Apple won't let you publish an app without paying them 30% of your total revenue. It's why Microsoft threatened to ban Steam (they later decided not to) so they could get all those games onto the Microsoft Store, and in response Valve ported gaming to Linux.
schainks 2 hours ago [-]
Wait until you find out what Google is paying Reddit…
tekacs 13 hours ago [-]
I mean that they're telling developers that they should use Cloudflare's platform to build agents, the kind of agents that would go across the web and act on behalf of users... but then they're also the ones blocking those requests.
This always engenders a solid amount of distaste from me, because much like Google and Chrome, it creates the incentive for you to treat yourself better than others. Especially coupled with the trust stuff. Of course, Cloudflare is always going to trust their own platform.
samrus 9 hours ago [-]
There are ways to use those agents ethically. Cloudflare isnt forcing you to be a dick with your agents. Infact the blocking side encourages you to use them ethically. I dont see the problem.
stingraycharles 11 hours ago [-]
I thought the issue people have with AI scrapers is the ones that DDoS sites to scrape everything for training purposes, rather than the ones that interactively query specific content to support an active conversation?
tekacs 10 hours ago [-]
If you read the article, you'll notice that they're explicitly making sure to block the ones that interactively query specific content too.
PunchyHamster 10 hours ago [-]
First one is problem because it brings no traffic back and occasionally DDoSes the site
Second one is problem because it DDoSes the site. "Just" queries for hundred thousand people (if you happened to be good source for that bit of knowledge) that don't bring actual people to your site is also a problem
inigyou 6 hours ago [-]
First one is not occasional, it is the anonymous global adversary probably ddosing your site right now. It never relents.
nirui 12 hours ago [-]
> not take down systems
"Block on pages with ads" is probably about preventing the AI crawlers from clicking on the ads which maybe considered cheating by the ad company.
If you want to prevent "bot attacks", maybe the "Block" option will do the trick.
But of course, to do all that you need to put some trust on Cloudflare, because they're the one identifying the bots from normal users.
For me, as someone who's hosting a Gitea instance behind Cloudflare, I have a Configuration Rule set that says: if the client is trying to access a URL that is beyond certain length limit, then trigger "Browser Integrity Check" and "I’m Under Attack", a.k.a stricter security checks.
The match expression of the rule looked something like this:
(
len(http.request.uri) > !!!!!SET LENGTH LIMIT!!!!! and
not lower(http.request.uri.path) contains ".git/" and
not lower(http.request.uri.path) contains "api/"
)
(The `!!!!!SET LENGTH LIMIT!!!!!` is an integer of the length limit you wanted to set)
This rule alone basically blocked all abusive bot traffic to almost zero for my site (https://i.imgur.com/LaOjjvV.png, see the traffic drop around 10 clock and Cloudflare mitigation kicks in).
But of course, you need to figure out your own rules based on the characteristic of the website. Also, you can be more creative: for example, my actual rule is more complex than that, it also checks to see if a cookie is not set, and only triggers when all condition are met:
(
len(http.request.uri) > !!!!!SET LENGTH LIMIT!!!!! and
not lower(http.request.uri.path) contains ".git/" and
not lower(http.request.uri.path) contains "api/" and
not http.cookie wildcard "*!!!!!COOKIE NAME!!!!!=!!!!!COOKIE VALUE!!!!!*"
)
then, as part two of that rule, I have a Response Header Transform Rules that says:
(
len(http.request.uri) <= !!!!!SET LENGTH LIMIT!!!!! and
not http.cookie contains "!!!!!COOKIE NAME!!!!!=!!!!!COOKIE VALUE!!!!!"
)
and if this Response Header Transform Rules is triggered, it sets the cookie `!!!!!COOKIE NAME!!!!!=!!!!!COOKIE VALUE!!!!!`.
(Note: `!!!!!COOKIE NAME!!!!!` and `!!!!!COOKIE VALUE!!!!!` are the variables you need to customize)
When you put the two rules together, it forces clients to access "shallow" (short URL) pages first as an user would normally do, before they can access "deeper" (long URL) content hosted on the site without triggering more strict security checks. If that makes sense.
Also, don't forget the cookie basically also dug a hole in the security setting. So it's really a balance between avoid annoying the user and protect your site. You need to be smart and be flexible about it, otherwise your users will just leave.
inigyou 8 hours ago [-]
> If you want to prevent "bot attacks", maybe the "Block" option will do the trick.
This morning, for the first time, I clicked on a link posted on HN and was told in no uncertain terms that I am a bot and will not be allowed to view this page. By Cloudflare.
Your strategy is fine and similar to one of the checks in go-away. The point is that the unknown global DDoS adversary is using a very simple scraper, which does not load images or CSS or scripts, does not set cookies, etc
15 hours ago [-]
colechristensen 13 hours ago [-]
I see Cloudflare as trying to forge the appropriate path ahead. Neither allowing the free-for-all nor trying to block everything isn't playing both sides, it's the path straight down the middle. Providing tools for producers and consumers to do things with permission and compensation.
jwr 7 hours ago [-]
I find it unsettling that we are willingly outsourcing the decision on who can access our sites to an increasingly dominant corporate entity.
The reasoning behind this is also flawed: blocking "bots" and "AI" means that our AI agents working for us are unable to do their work for us, because of knee-jerk bot-blocks.
fc417fc802 15 hours ago [-]
Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go this route.
prologic 13 hours ago [-]
PoW schemes like Anubis don't work. Increasingly bots are using headless browsers and are basically able to solve captchas, proof-of-work(s) and basically bypass all any any attempts to block them. It's becoming impossible to stop bots from hammering your sites/services for unwanted traffic.
inigyou 8 hours ago [-]
There's one or a small number of actors who are DDoSing the internet right now, and those are the ones you care about. What's the purpose for blocking non-DDoS bots?
fc417fc802 11 hours ago [-]
You are asserting that they don't work without explanation or evidence. Meanwhile it isn't clear why they wouldn't and indeed they appear to accomplish the stated goal of severely rate limiting scrapers.
codethief 5 hours ago [-]
IIRC there was a blog post here on HN recently explaining in detail why Anubis doesn't keep out AI bots.
prologic 11 hours ago [-]
Surely I don't need to explain how headless browsers work, bot proxies and distributed crawlers to you do I? If you haven't experienced this first-hand, I can understand. Maybe it warrants a blog post. But I can assure you, these don't work at scale, they might only stop the "less sophisticated" bots and maybe (just maybe) some unwanted spam.
fc417fc802 5 hours ago [-]
What do the browsers being headless or the traffic proxied have to do with anything? Each client still has to solve a PoW challenge regardless. That mitigates the (possibly unintentional) DoS by making it exorbitantly expensive.
gib444 11 hours ago [-]
> You are asserting that they don't work without explanation or evidence.
Oh, no, that isn't how it works — you made a claim first (albeit indirectly), with no evidence. It's on you
inigyou 6 hours ago [-]
Evidence: they solved many people's problems
fc417fc802 5 hours ago [-]
That's not how it works. There are cases where established norms or common sense dictate a certain default assumption. This is one of them.
Still, I'll humor your absurd request by reminding you of the many success stories that have repeatedly made the front page of HN.
Beyond that we have cryptocurrencies. If you have knowledge of a generalized solution for defeating PoW schemes then why are you posting here instead of making yourself a billionaire?
m00dy 12 hours ago [-]
The most recent captchas that involve basic reasoning can block majority of bots today.
prologic 12 hours ago [-]
Name one, that an LLM can't solve. The only real mechanism that continues to still work (but doesn't scale) is "Human in the loop".
timpera 6 hours ago [-]
Please don't use Anubis, it makes visiting websites very difficult (often multiple minutes wait times) on old hardware and low-end smartphones.
altairprime 5 hours ago [-]
As things stand right now, the current web may not be possible to maintain for old and low-end phones given the costs imposed by AI training crawlers.
Please propose an alternative to both Cloudflare and Anubis, that shields websites against inhuman traffic without frequent operator intervention (or otherwise negates the capacity costs
they pay for AI crawling) and is compatible with low-end smartphones.
Certainly, I imagine Anubis would be interested in adopting it if it’s effective!
fc417fc802 5 hours ago [-]
Personally I think that's a better outcome than the alternative.
However another option is to support both. Have a challenge page that requires the visitor to select one of several options. Cloudflare could be one of those.
gobip 5 hours ago [-]
What's the alternative? We're stuck, either cloudflare or anubis
r_lee 4 hours ago [-]
what kind of hardware would take minutes to complete the checks?
gruez 14 hours ago [-]
>I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites.
???
Unless you're browsing around with the googlebot user agent string, you should be getting turnstile challanges at most, not blocks. And if you're getting a turnstile challenge it's unclear how it's different than an anubis challenge. If you're outright blocked, it's probably a site decision (eg. block all VPNs or block everyone not from a given country) rather than cloudflare's.
amatecha 14 hours ago [-]
Try using Firefox with "resist fingerprinting" enabled, on OpenBSD, if you want an idea of how insidious Cloudflare's increasing prevalence is for the openness of the web. Inexplicable 403 responses for entire domains for no apparent reason, no "captcha", just blocked outright. Heck, it happens on Linux too, and even without "resist fingerprinting" enabled.
ipaddr 14 hours ago [-]
Happens with older chrome browsers as well. Last supported windows 7 version. Started this month.
gucci-on-fleek 14 hours ago [-]
> And if you're getting a turnstile challenge it's unclear how it's different than an anubis challenge.
You are guaranteed to pass an Anubis challenge eventually [0], whereas it's possible to get stuck forever in an infinitely-looping Turnstile challenge.
> If you're outright blocked, it's probably a site decision (eg. block all VPNs or block everyone not from a given country) rather than cloudflare's.
Cloudflare blocks legitimate users itself sometimes [1].
[0]: Unless you run into a bug, but Anubis is open source, so you can always submit a patch upstream. I've done this myself, and I can confirm that it's relatively straightforward.
>You are guaranteed to pass an Anubis challenge eventually [0]
That's a double edged sword because bots will eventually get through too, and unlike humans, their time is dirt cheap.
>Cloudflare blocks legitimate users itself sometimes [1].
I never ran into this issue despite using seemingly maximally suspicious configs like tor browser. I can't say the same for some other vendors.
gucci-on-fleek 13 hours ago [-]
> That's a double edged sword because bots will eventually get through too, and unlike humans, their time is dirt cheap.
Yeah, I really have no idea why Anubis works right now: residential proxies are far more expensive than compute, yet the bots seem to have no problem obtaining millions of residential IPs, but they give up on even short-ish Anubis challenges.
> I never ran into this issue despite using seemingly maximally suspicious configs like tor browser. I can't say the same for some other vendors.
Yeah, I don't like the Cloudflare challenges, but in the past 5 years I've only had it outright block me once, and that fixed itself after 15 minutes. And I use Firefox on Linux with various privacy extensions, so my browser probably appears at least moderately suspicious.
Whereas I've been trapped in impossible ReCaptcha loops quite a few times, which is still better than vague error messages that magically go away when I switch to something not running Linux. So I'll begrudgingly accept that Turnstile is the least user-hostile product on the market right now.
inigyou 8 hours ago [-]
Because the (singular) global ddos adversary is using a very dumb scraper, even dumber than wget --mirror
raincole 13 hours ago [-]
The only reason Anubis works right now is that it's not very commonly used and the bots are not optimized to bypass it (yet).
inigyou 8 hours ago [-]
The bots that are problems right now are not the ones that run Anubis challenges. There is no universal anti-bot, only ones that work at certain times.
fc417fc802 11 hours ago [-]
> unlike humans, their time is dirt cheap.
On the contrary I only want to visit a few pages on a given site and have an entire laptop at my disposal. Meanwhile for bots efficiency is key. A serious scraper (ie the type of actor that actually causes material problems for site operators) is performing tens or hundreds of pages loads per second per core spread across thousands of sites. Making a single page load take even half a second of cpu time is a massive win for the site operator.
benhurmarcel 3 hours ago [-]
I’ve been blocked by Turnstile without possibility of challenge, no way to access nor know why. And it was with standard iOS and Windows machines and no VPN or anything. It happens.
inigyou 8 hours ago [-]
Yesterday Cloudflare told me I was a bot and couldn't access this.
m00dy 12 hours ago [-]
yeah, googlebot user agent thing was a thing maybe around 10 years ago but I don't think it's working today. Maybe if you start the http session from gcloud, it might still work.
matheusmoreira 14 hours ago [-]
> Please consider installing one of the many PoW schemes such as anubis
Why not go all the way and mine monero instead of just completely wasting the work?
akersten 11 hours ago [-]
because then someone will complain that you're stealing their CPU cycles or something, on the website they chose to visit and execute
matheusmoreira 14 minutes ago [-]
"Stealing" CPU cycles is exactly what PoW bot protection is doing. The whole point is to add cost to the bots so they decide it's too expensive and give up.
The only difference is the cycles are getting converted into heat now. They could be getting converted into monero instead. It's still heat but at least creator got some money for it.
neya 13 hours ago [-]
Adding the link to GitHub here if anyone is curious:
Not sure why Anubis is getting so much hype on HN, but honestly, it is not the solution. A real solution would use behavioral modeling. Most browser fingerprinting issues are already largely solved anyway.
bornfreddy 8 hours ago [-]
As someone who is regularly blocked by CF and G - no, they are not. Even worse, they should not be, because that would mean the complete end of privacy online (not that we are far from that). Anubis works because at scale it wastes bots' resources (time mainly).
inigyou 8 hours ago [-]
it works because the global ddos adversary doesn't run javascript
m00dy 8 hours ago [-]
I think you’re confusing bots that crawl a website with bots used to launch DDoS attacks.
fc417fc802 11 hours ago [-]
PoW is a reasonable solution as a fallback when other metrics flag a client.
microtonal 5 hours ago [-]
How? First, they can solve the Anubis challenge with native code, so they can solve them faster than genuine users. Second, the cost is nothing compared to training LLMs, plus they will just move the work to the residential proxies that they have access to (so, someone is paying through their TV's electricity bill).
Anubis only work(s|ed) great for a while when crawlers were not prepared for these challenges. Security through obscurity.
fc417fc802 4 hours ago [-]
> they will just move the work to the residential proxies that they have access to
The proxies I am familiar with do not offer arbitrary code execution. I think you're thinking of a botnet.
Regarding native code, the current crop of solutions seem to work well enough for now. Ultimately a challenge response protocol should be standardized and browsers should ship a native implementation. In the meantime WASM likely gets you close enough to native.
microtonal 4 hours ago [-]
Regarding native code, the current crop of solutions seem to work well enough for now.
I just quoted a toot in another submission, adding it here since it is relevant:
We apologize for a period of extreme slowness today. The army of AI crawlers just leveled up and hit us very badly. [...] It seems like the AI crawlers learned how to solve the Anubis challenges. [...] However, we can confirm that at least Huawei networks now send the challenge responses and they actually do seem to take a few seconds to actually compute the answers. It looks plausible, so we assume that AI crawlers leveled up their computing power to emulate more of real browser behaviour to bypass the diversity of challenges that platform enabled to avoid the bot army.
So is it possible to say "No bots except Google, OpenAI, Grok, Claude and Perplexity"?
As far as I can tell, Google is the only one sending me visitors. And the other big AI players might do so in the future.
Another option would be "No anonymous bots". So at least if a bot would want to crawl my site, they would have to identify themselves. Since the rise of the AI bots, I am getting hurt badly with insane amounts of requests from residential IPs that mimic real humans. The only difference being they don't make me any money. Only produce costs.
By the way, how is the situation over at Amazon's Cloudfront? Do they offer something that helps? Anyone here with them?
nicbou 9 hours ago [-]
Google is actively working on not sending you visitors anymore.
holografix 14 hours ago [-]
What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl.
What would force their hand?
It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc.
That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.
jerf 13 hours ago [-]
I think the best answer is, nobody knows. The previous equilibrium for content scraping for search engines on the internet was already at times an uncomfortable one. But I agree that from a game theory perspective, "the AI bots take and give nothing back in return" is not just hyperbole, it's the actual situation. If Google is successful in what seems to be its plans and it becomes a box where you type a question and Google gives you an answer and only a vanishing fraction of the users click through to any underlying website, that instantly eliminates the entire value proposition for vast swathes of the web to actually be on the web.
Something has to happen or Google will end up starved and locked out of everything, by means both technical and legal. Then nobody gets anything.
I don't have the answer as to what happens next, and I doubt anyone else who proclaims one super confidently. But we can do some constraints analysis. There is no world where everyone works for free so Google and other AI engines can get all the value from the content, so we can eliminate those possibilities. I think we can safely discard the world(s) in which all content production just stops. However, off the top of my head, it's hard to get much tighter than that, and that definitely leaves a world where effectively everything everywhere ends up going pay-to-access.
Microtransactions have, to date, failed comprehensively, though, so the constraints on what "everything is pay-to-access" gets weird without them.
And there is never guarantee that there is any solution to any set of constraints. Things can end up overconstrained in reality as easily as a math problem. I don't actually think it'll go that way, but when analyzing this question I think it's important to not let "but $SOMETHING just has to have some way to work, because... uh... it has to!" Let the constraints do the talking. You could end up with a scenario where all content of any value is locked down, and it's fundamentally difficult and expensive to ever access or discover it, and consequently the entire content production industry radically contracts compared to its current size, if there is no pragmatic solution to microtransactions that is low-enough friction to get over the psychological and economic hurdles that have killed it to date. If everything is locked behind "macrotransactions" that's a much smaller commercial web. Probably a much higher quality one, too, but at a pretty stiff cost.
inigyou 8 hours ago [-]
The internet becomes full of free propaganda since it's not the consumer who pays for that?
jerf 3 hours ago [-]
That's the current state of the internet. In this world it wouldn't be "the Internet" full of propaganda, it would be the AI search engines the propaganda would get concentrated into. For "national security", don't you know. And it would be much easier for them for having an even smaller target. At least the current internet lets you cross-check one source of free propaganda against another today. (Although whether the truth is in any meaningful way "between" any set of them is another question.)
One possible scenario is that that is simply it for the internet as an information source; the search engine's AIs get captured and there becomes effectively no way to discover any of the content already on there.
But then again, people will react to that and do something. Kagi would grow and others too. The more interesting question is whether the governments that captured Google's AI would let them or if suddenly it would be discovered that copyright law doesn't permit search engines to do that. Would that be inconsistent with letting the Approved AIs access whatever they want and chew on it even harder? Yeah, and they wouldn't care.
PeterStuer 9 hours ago [-]
Universal tax collector of the internet. A penny for every page access. ADOG will not mind as it cements their incumbent status and pulls up the drawbridge by erecting a huge financial barrier for any new entrant.
sandeepkd 14 hours ago [-]
I get a mixed feeling about all this. Cloudflare is unilaterally making all these decisions which impact the whole internet traffic flow. Taking the lead is one thing, however decisions like this should have the direct involvement of Internet Engineering Task Force (IETF) to account for all stakeholders, otherwise we run into the situation of a fragmented internet
usef- 14 hours ago [-]
I'm curious if this is just to pressure Google into separating their crawlers
sandeepkd 14 hours ago [-]
This is more of a way to create a unauthorized toll tax on highway. There is a problem indeed, however if the proposed direction by cloud flare is to solve it or benefit out of it is a debatable topic.
inigyou 6 hours ago [-]
That wouldn't make any sense from anyone's perspective. They would just pretend to separate them and not.
colechristensen 13 hours ago [-]
Are they unilateral? Maybe the defaults? For anyone competent they're settings freely chosen.
siva7 7 hours ago [-]
"I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl."
You seem to misunderstand. You pay or you die. There's nothing in between. Cloudflare will happily collect the tax. As does Apple (collecting 20B$ yearly from Google for the "tax"). Cloudflare's users also won't mind about how the company handles ADOG as long as they get a chunk of the cake by getting freebies and cheap services.
qntmfred 3 hours ago [-]
*AGOD
deadbabe 13 hours ago [-]
Pay to crawl is already here.
The usage patterns of how people pay and use AI is basically the same model the web should be using: you pay a small bit of money to access monetized pages, just how you pay a small bit of money to get AI responses.
It just needs people and browsers to get onboard with protocols. Crawlers will have no choice but to pay for content behind these 402 gateways.
guyn 6 hours ago [-]
I tried this, blocking AI training blocked the Google search bots and cut my traffic in half. I would not recommend.
graeme 16 hours ago [-]
Has there been any update on the pay per crawl program?
noduerme 11 hours ago [-]
I wonder if this has anything to do with the cf bug that stripped all POST data from requests to a SPA I manage for 4-5 hours last week. That was a real good time, figuring out that it wasn't trying to show challenges or anything. Default setting for any web app protection from cloudflare should always be "off" unless you're under attack, and then who knows what settings will or won't break your configuration.
arjie 15 hours ago [-]
This is fine so long as it’s easy for me to turn off. I just don’t want to accidentally lose all AI traffic one day.
PeterStuer 9 hours ago [-]
Most of the internet unfortunatly ploinks fown a 'free' service and never even looks at what it defaults to. E.g., quite a lott of cloudflare "protected" sites block their rss feeds from being read by machine.
inigyou 8 hours ago [-]
This is deliberate from Cloudflare. It wants to block Google's access to as much of the internet as possible.
dmortin 7 hours ago [-]
Do you actually get any signifcant AI traffic today?
timpera 6 hours ago [-]
Blocking bots should not be the default behavior, it should be opt-in. Is Cloudflare trying to play both sides in order to secure a fee on every pay-to-crawl in the future?
inigyou 6 hours ago [-]
Yes. Obviously.
pluc 5 hours ago [-]
Cloudflare's business model is playing both sides. They sell anonymity to attackers and mitigation/protection to the victims.
Fizz43 12 hours ago [-]
>So, instead of defining a bot primarily as “AI” or not, our updated approach to classification will ask deeper questions about bot or agent behavior: What are they doing on my site? What are they storing? And how will they reshare my content?
I dont get this. The question is are they a bot or a human. It doesnt matter what they are doing I dont want bots on my site.
akersten 11 hours ago [-]
> I dont get this. The question is are they a bot or a human. It doesnt matter what they are doing I dont want bots on my site.
Do you want your site to be discoverable by a search engine? (How do you think that occurs?)
Terr_ 9 hours ago [-]
Let's also spare a moment for accessibility issues. For example, is it that wrong for a blind person to invoke a tool that describes a picture when there's no alt-text? Or something which describes/transcribes audio for the deaf?
Where do we draw the line between a personal-bot and a custom browser?
Imustaskforhelp 8 hours ago [-]
Could there be a proper committee/association which can have all good faith search engines (which work for the purpose of search rather than AI related) which could show their proper IP/networks and cloudflare could have an option which supports all search but not any AI botnets.
That being said, the issue with it right now is that Google uses the same crawler for both AI training and search, so this step puts a small pressure on google nonetheless to hopefully split them.
inigyou 8 hours ago [-]
Why not. You need a reason. A browser is just a bot that renders pages.
zzzeek 15 hours ago [-]
this is annoying, it makes a big deal about "Back when we announced pay-per-crawl"...
I want pay-per-crawl. I clicked the link for it a year ago, got presented with a "request access" button, I "requested access" and obviously since I'm nobody I heard absolutely nothing. Now they're touting the link again, I checked, still that same "request access" button. I have no idea if anyone even has access to this feature.
I don't care about all this other stuff, I want the AI crawlers to pay me cash. Because boy do those fuckers want to crawl me. I'll gladly double the size of my gerrit/jenkins servers to keep up with the load if these stupid bots want to pay to crawl every jenkins build artifact and every changeset source file on the server, as they really seem to want to do.
sneak 11 hours ago [-]
Website operators don’t lose anything when people download the content from their website and use it.
There is no technical mechanism whereby it is actually possible to allow people to read your webpage and not use it for other things. You can’t give responses that say “this is ok for indexing but not for training”. Anyone trying to sell you this sort of technology is lying.
edifierxuhao 6 hours ago [-]
[flagged]
fllkfsalkdsfds 1 hours ago [-]
[dead]
julian-vix 5 hours ago [-]
[flagged]
youre-wrong3 15 hours ago [-]
[dead]
paul7986 14 hours ago [-]
[dead]
ray_v 16 hours ago [-]
So, in summary: still the honors system. Got it. thanks.
zx8080 16 hours ago [-]
What's the "honors system"?
willy_k 14 hours ago [-]
An honor system is a system without any (explicit) external enforcement of rules. For example an unnattended fruit stand, where people are trusted to be honorable and leave payment for the fruit they take.
> Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).
That's exactly what antitrust laws are supposed to do, and I hope at least EU regulators take action. Every single Googlebot crawl in your access logs is a trace for damages.
> Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search. https://developers.google.com/crawling/docs/crawlers-fetcher...
Exclusion from grounding does mean that your site won't get sourced in the AI overview, but I'm not sure what the click through rates are like on those.
Saying this as a European who is pro EU.
Why should they?
How is that not conflict of interest??
Huh?
Search and AI are hand-in-hand.
They both rely on embeddings. (Unless you still do keyword-only search, but that's not as good.)
that’s just my take/feedback, take it or leave it. i won’t be engaging further as i already feel i’m going against the site guidelines with this!
This is unfortunately nothing new. There's no correct way to tell them to fuck off, they do not care, and they never will do. People have even taken them to court over this
Also I should note there are lots of fake Googlebots...
Akamai, Fastly, etc only take big customers who know what they're doing. You need to sign a proper contract with them. They aren't low-friction.
It's kind of exhausting seeing Cloudflare playing both sides of the arms race.
I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this.
> This also lines up the incentive model we want to foster. Losing trusted status across the more than 20% of web domains that sit behind Cloudflare is a deterrent with teeth. Trust becomes something you can carry with you, and something you can lose.
And even more so, LLM language aside, fun and fascinating to see them flagrantly calling out their position here as if it's a positive.
This always engenders a solid amount of distaste from me, because much like Google and Chrome, it creates the incentive for you to treat yourself better than others. Especially coupled with the trust stuff. Of course, Cloudflare is always going to trust their own platform.
Second one is problem because it DDoSes the site. "Just" queries for hundred thousand people (if you happened to be good source for that bit of knowledge) that don't bring actual people to your site is also a problem
"Block on pages with ads" is probably about preventing the AI crawlers from clicking on the ads which maybe considered cheating by the ad company.
If you want to prevent "bot attacks", maybe the "Block" option will do the trick.
But of course, to do all that you need to put some trust on Cloudflare, because they're the one identifying the bots from normal users.
For me, as someone who's hosting a Gitea instance behind Cloudflare, I have a Configuration Rule set that says: if the client is trying to access a URL that is beyond certain length limit, then trigger "Browser Integrity Check" and "I’m Under Attack", a.k.a stricter security checks.
The match expression of the rule looked something like this:
(The `!!!!!SET LENGTH LIMIT!!!!!` is an integer of the length limit you wanted to set)This rule alone basically blocked all abusive bot traffic to almost zero for my site (https://i.imgur.com/LaOjjvV.png, see the traffic drop around 10 clock and Cloudflare mitigation kicks in).
But of course, you need to figure out your own rules based on the characteristic of the website. Also, you can be more creative: for example, my actual rule is more complex than that, it also checks to see if a cookie is not set, and only triggers when all condition are met:
then, as part two of that rule, I have a Response Header Transform Rules that says: and if this Response Header Transform Rules is triggered, it sets the cookie `!!!!!COOKIE NAME!!!!!=!!!!!COOKIE VALUE!!!!!`.(Note: `!!!!!COOKIE NAME!!!!!` and `!!!!!COOKIE VALUE!!!!!` are the variables you need to customize)
When you put the two rules together, it forces clients to access "shallow" (short URL) pages first as an user would normally do, before they can access "deeper" (long URL) content hosted on the site without triggering more strict security checks. If that makes sense.
Also, don't forget the cookie basically also dug a hole in the security setting. So it's really a balance between avoid annoying the user and protect your site. You need to be smart and be flexible about it, otherwise your users will just leave.
This morning, for the first time, I clicked on a link posted on HN and was told in no uncertain terms that I am a bot and will not be allowed to view this page. By Cloudflare.
Your strategy is fine and similar to one of the checks in go-away. The point is that the unknown global DDoS adversary is using a very simple scraper, which does not load images or CSS or scripts, does not set cookies, etc
The reasoning behind this is also flawed: blocking "bots" and "AI" means that our AI agents working for us are unable to do their work for us, because of knee-jerk bot-blocks.
Oh, no, that isn't how it works — you made a claim first (albeit indirectly), with no evidence. It's on you
Still, I'll humor your absurd request by reminding you of the many success stories that have repeatedly made the front page of HN.
Beyond that we have cryptocurrencies. If you have knowledge of a generalized solution for defeating PoW schemes then why are you posting here instead of making yourself a billionaire?
Please propose an alternative to both Cloudflare and Anubis, that shields websites against inhuman traffic without frequent operator intervention (or otherwise negates the capacity costs they pay for AI crawling) and is compatible with low-end smartphones.
Certainly, I imagine Anubis would be interested in adopting it if it’s effective!
However another option is to support both. Have a challenge page that requires the visitor to select one of several options. Cloudflare could be one of those.
???
Unless you're browsing around with the googlebot user agent string, you should be getting turnstile challanges at most, not blocks. And if you're getting a turnstile challenge it's unclear how it's different than an anubis challenge. If you're outright blocked, it's probably a site decision (eg. block all VPNs or block everyone not from a given country) rather than cloudflare's.
You are guaranteed to pass an Anubis challenge eventually [0], whereas it's possible to get stuck forever in an infinitely-looping Turnstile challenge.
> If you're outright blocked, it's probably a site decision (eg. block all VPNs or block everyone not from a given country) rather than cloudflare's.
Cloudflare blocks legitimate users itself sometimes [1].
[0]: Unless you run into a bug, but Anubis is open source, so you can always submit a patch upstream. I've done this myself, and I can confirm that it's relatively straightforward.
[1]: https://news.ycombinator.com/item?id=43329320
That's a double edged sword because bots will eventually get through too, and unlike humans, their time is dirt cheap.
>Cloudflare blocks legitimate users itself sometimes [1].
I never ran into this issue despite using seemingly maximally suspicious configs like tor browser. I can't say the same for some other vendors.
Yeah, I really have no idea why Anubis works right now: residential proxies are far more expensive than compute, yet the bots seem to have no problem obtaining millions of residential IPs, but they give up on even short-ish Anubis challenges.
> I never ran into this issue despite using seemingly maximally suspicious configs like tor browser. I can't say the same for some other vendors.
Yeah, I don't like the Cloudflare challenges, but in the past 5 years I've only had it outright block me once, and that fixed itself after 15 minutes. And I use Firefox on Linux with various privacy extensions, so my browser probably appears at least moderately suspicious.
Whereas I've been trapped in impossible ReCaptcha loops quite a few times, which is still better than vague error messages that magically go away when I switch to something not running Linux. So I'll begrudgingly accept that Turnstile is the least user-hostile product on the market right now.
On the contrary I only want to visit a few pages on a given site and have an entire laptop at my disposal. Meanwhile for bots efficiency is key. A serious scraper (ie the type of actor that actually causes material problems for site operators) is performing tens or hundreds of pages loads per second per core spread across thousands of sites. Making a single page load take even half a second of cpu time is a massive win for the site operator.
Why not go all the way and mine monero instead of just completely wasting the work?
The only difference is the cycles are getting converted into heat now. They could be getting converted into monero instead. It's still heat but at least creator got some money for it.
https://github.com/techaroHQ/anubis
Anubis only work(s|ed) great for a while when crawlers were not prepared for these challenges. Security through obscurity.
The proxies I am familiar with do not offer arbitrary code execution. I think you're thinking of a botnet.
Regarding native code, the current crop of solutions seem to work well enough for now. Ultimately a challenge response protocol should be standardized and browsers should ship a native implementation. In the meantime WASM likely gets you close enough to native.
I just quoted a toot in another submission, adding it here since it is relevant:
We apologize for a period of extreme slowness today. The army of AI crawlers just leveled up and hit us very badly. [...] It seems like the AI crawlers learned how to solve the Anubis challenges. [...] However, we can confirm that at least Huawei networks now send the challenge responses and they actually do seem to take a few seconds to actually compute the answers. It looks plausible, so we assume that AI crawlers leveled up their computing power to emulate more of real browser behaviour to bypass the diversity of challenges that platform enabled to avoid the bot army.
https://social.anoxinon.de/@Codeberg/115033790447125787
As far as I can tell, Google is the only one sending me visitors. And the other big AI players might do so in the future.
Another option would be "No anonymous bots". So at least if a bot would want to crawl my site, they would have to identify themselves. Since the rise of the AI bots, I am getting hurt badly with insane amounts of requests from residential IPs that mimic real humans. The only difference being they don't make me any money. Only produce costs.
By the way, how is the situation over at Amazon's Cloudfront? Do they offer something that helps? Anyone here with them?
What would force their hand?
It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc.
That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.
Something has to happen or Google will end up starved and locked out of everything, by means both technical and legal. Then nobody gets anything.
I don't have the answer as to what happens next, and I doubt anyone else who proclaims one super confidently. But we can do some constraints analysis. There is no world where everyone works for free so Google and other AI engines can get all the value from the content, so we can eliminate those possibilities. I think we can safely discard the world(s) in which all content production just stops. However, off the top of my head, it's hard to get much tighter than that, and that definitely leaves a world where effectively everything everywhere ends up going pay-to-access.
Microtransactions have, to date, failed comprehensively, though, so the constraints on what "everything is pay-to-access" gets weird without them.
And there is never guarantee that there is any solution to any set of constraints. Things can end up overconstrained in reality as easily as a math problem. I don't actually think it'll go that way, but when analyzing this question I think it's important to not let "but $SOMETHING just has to have some way to work, because... uh... it has to!" Let the constraints do the talking. You could end up with a scenario where all content of any value is locked down, and it's fundamentally difficult and expensive to ever access or discover it, and consequently the entire content production industry radically contracts compared to its current size, if there is no pragmatic solution to microtransactions that is low-enough friction to get over the psychological and economic hurdles that have killed it to date. If everything is locked behind "macrotransactions" that's a much smaller commercial web. Probably a much higher quality one, too, but at a pretty stiff cost.
One possible scenario is that that is simply it for the internet as an information source; the search engine's AIs get captured and there becomes effectively no way to discover any of the content already on there.
But then again, people will react to that and do something. Kagi would grow and others too. The more interesting question is whether the governments that captured Google's AI would let them or if suddenly it would be discovered that copyright law doesn't permit search engines to do that. Would that be inconsistent with letting the Approved AIs access whatever they want and chew on it even harder? Yeah, and they wouldn't care.
You seem to misunderstand. You pay or you die. There's nothing in between. Cloudflare will happily collect the tax. As does Apple (collecting 20B$ yearly from Google for the "tax"). Cloudflare's users also won't mind about how the company handles ADOG as long as they get a chunk of the cake by getting freebies and cheap services.
The usage patterns of how people pay and use AI is basically the same model the web should be using: you pay a small bit of money to access monetized pages, just how you pay a small bit of money to get AI responses.
It just needs people and browsers to get onboard with protocols. Crawlers will have no choice but to pay for content behind these 402 gateways.
I dont get this. The question is are they a bot or a human. It doesnt matter what they are doing I dont want bots on my site.
Do you want your site to be discoverable by a search engine? (How do you think that occurs?)
Where do we draw the line between a personal-bot and a custom browser?
That being said, the issue with it right now is that Google uses the same crawler for both AI training and search, so this step puts a small pressure on google nonetheless to hopefully split them.
I want pay-per-crawl. I clicked the link for it a year ago, got presented with a "request access" button, I "requested access" and obviously since I'm nobody I heard absolutely nothing. Now they're touting the link again, I checked, still that same "request access" button. I have no idea if anyone even has access to this feature.
I don't care about all this other stuff, I want the AI crawlers to pay me cash. Because boy do those fuckers want to crawl me. I'll gladly double the size of my gerrit/jenkins servers to keep up with the load if these stupid bots want to pay to crawl every jenkins build artifact and every changeset source file on the server, as they really seem to want to do.
There is no technical mechanism whereby it is actually possible to allow people to read your webpage and not use it for other things. You can’t give responses that say “this is ok for indexing but not for training”. Anyone trying to sell you this sort of technology is lying.