An update on Wayback Machine access (blog.archive.org)

518 points by ChrisArchitect 15 hours ago

simonw 14 hours ago

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

Kodiack 9 hours ago

I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests.

However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a small donation their way. They provide an incredibly valuable service and I love the benefit that I get from them just for personal side projects.

sippingabonedry 7 hours ago

How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks?

Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...

Kodiack 6 hours ago

fc417fc802 7 hours ago

4thguy an hour ago

Kudos on that. I don't know what site you're hosting, but I appreciate knowing that it is there

mkatx 8 hours ago

This is the way to go! Cut the cat and mouse, win win ish.

packetslave 14 hours ago

This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.

bsimpson 14 hours ago

It's an open secret that you can often circumvent paywalls by searching Wayback.

koolala 5 hours ago

zymhan 9 hours ago

gambiting 14 hours ago

pantsforbirds 13 hours ago

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!

Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.

subarctic 13 hours ago

What if they charged money? Is it something you'd pay for?

bonestamp2 13 hours ago

msephton 13 hours ago

pantsforbirds 4 hours ago

bee_rider 12 hours ago

bradly 14 hours ago

Just yesterday from my one of my sessions with Sol:

> Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

TeMPOraL 13 hours ago

As it should.

Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

bradly 13 hours ago

matt_heimer 10 hours ago

I wonder if the entire internet is going to slowly move behind logins and allow lists for specific trusted crawlers at some point.

Open access doesn't seem sustainable.

But I might just grumpy about spending another hour this week adjusting rules to prevent bots.

intrasight 9 hours ago

Rather than logins or regional filters, how about they just be a content provider to local libraries and perhaps use an app like Libby.

autoexec 11 hours ago

I've personally been using the Wayback Machine more often because I increasingly find myself being blocked from websites who are trying to keep out scrapers even though I'm just a regular person with JS disabled (along with a bunch of other stuff)

eek2121 10 hours ago

Sites are getting too overzealous with blocking IMO. I got blocked for several hours by huggingface simply because my download didn't complete and I had to retry. It gave me error 429, suggested I login, and the login page wouldn't load because error 429.

A popular tech news site blocked my phone because of Apple Private Relay. That didn't last long because their traffic fell off a cliff when that happened.

Many sites are throwing more captchas at the problem, without understanding that captchas don't actually help with LLMs, they just hinder normal users and primitive scripts. LLMs solve captchas just fine.

Some big sites have put up improved paywalls. I'm fine with subscribing to a quality site, however, WSJ and all the other big media sites routinely spit out regurgitated garbage that can be had for free elsewhere (and due to political spin, their garbage is less valuable than the free versions of said content).

Some folks are declaring the internet dead. I wouldn't go that far, however, I will say that a reckoning is going to happen, especially when advertisers figure out that most ads served on basically every website are no longer viewed by humans.

simonjgreen 2 hours ago

Nearly every time a link is posted to HN to a site behind some form of wall, a high voted comment on the post will be a link to an archive site bypassing the owners wall. Bot owners are not the only ones routinely circumventing the choices of content owners.

RobotToaster 13 hours ago

Do they offer bulk torrent downloads as an alternative?

QuantumNomad_ 12 hours ago

Once upon a time some people explored backing up the Internet Archive.

However, that experiment ended. They mention there were some learnings and they then say:

> The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.

https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK

I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive.

I know that individual items have torrents. And I’ve downloaded a few that way but always it ends up only using the “web seed” (i.e. the BitTorrent client is retrieving the files from IA via HTTP) because there are no one seeding some random single item I found. Plus, those torrents are unreliable sometimes because they include meta data files that were since updated but the torrent was not updated and so the web seed is giving the updated files that don’t match what the torrent says their hashes should be. So then you have to jump through some extra hoops to fix that and then resume the download, and all the while the HTTP connections to IA servers time out because their servers are overloaded. So when I say I wonder about possibilities of using BitTorrent I mean to retrieve whole collections of many items instead of individual ones, and with actual other peers instead of just having it put load on IA HTTP servers.

alightsoul 10 hours ago

giantrobot 10 hours ago

echelon 13 hours ago

I would love to be able to download every page of a given domain as an archive, and I'd pay to do this.

msephton 13 hours ago

petcat 13 hours ago

carlosjobim 12 hours ago

ezekiel68 9 hours ago

You might be right but -- why would they have watied until these recent weeks?

luckylion 14 hours ago

What sites would they be targeting? Generic "just give me anything"? Whenever I check regular sites on IA, the coverage is spotty -- they'll have the homepage and a few important pages, but it quickly fizzles out.

Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.

hedora 7 hours ago

My use of wayback has skyrocketed recently due to anti-bot measures.

I often cannot get past captchas, and archive.org is one of the fallbacks I try.

However, archive.is, etc are more reliable.

I wish the internet archive acted more like a library system, where multiple organizations could mirror the content.

They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives down yet.

jader201 13 hours ago

> I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression.

> we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator.

Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.

throwawayk7h 6 hours ago

Perhaps it would be sensible for the wayback machine to not show paywalled articles for the first, say, 3 months.

toomuchtodo 14 hours ago

It is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection. Hard to defend against abuse of a public resource when its intent is public access with as little restriction as possible.

https://en.wikipedia.org/wiki/Tragedy_of_the_commons

(no affiliation)

ronsor 14 hours ago

Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy.

On the other hand, the Internet Archive is a non-profit offering a free public resource.

toomuchtodo 14 hours ago

TZubiri 9 hours ago

Nah, there's actual value in hitting historical versions and with agents the gap between "how long has this product been offered by this company" and "I should go to wayback machine and do a binary search to find the earliest snapshot that contains this product offering " has closed.

e40 7 hours ago

I say name and shame!

basilikum 13 hours ago

Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.

The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.

If you got some money to spare, consider donating to them. They need it.

ternaryoperator 12 hours ago

I donate to them every year b/c I fully agree they’re doing a thankless critical job very well.

j79 11 hours ago

Thank you for the inspiration! I just made my first donation.

niuzeta 9 hours ago

I've been donating $5 to them monthly for I don't know how long. I've only recently bumped it up to $25. They're the heroes of the internet age

superxpro12 13 hours ago

fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future.

The future is bleak :\

mrguyorama 12 hours ago

I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side.

It's going to suddenly be extremely valuable that wikipedia didn't settle for having a small rainy day fund and instead ceaselessly grabbed every fucking donation they could for two decades so they can fight such a legal battle.

jasonfarnon 9 hours ago

pvab3 11 hours ago

nephihaha 11 hours ago

userbinator 4 hours ago

The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic

Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.

I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.

robotmay 10 hours ago

Unrelated, but this week I've been on a memory binge with the Wayback Machine, trying to find old content of mine from the early 2000s. Took me a while but I've finally put together a good bit of info about myself at the time that I'd completely forgotten, and it's all thanks to the Internet Archive storing my little gaming review website from when I was 16. I could barely remember any of the other stuff, it's been genuinely surprising figuring out what I'd forgotten. I couldn't even remember most domains I owned aside from one, which I used as the starting point.

Still can't remember what my Tripod site address was, but that might be lost to time.

Thank you, Archive.org.

BeetleB 14 hours ago

Wow, but I wonder if there's more to it.

I've not been able to access web.archive.org from my work computer - I always get the 429 error.

But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

flexagoon 14 hours ago

I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP

hedora 7 hours ago

Are there any decent/reputable residential proxy companies? I’m pretty sure I’ll end up needing one occasionally, for those days when my residential IP has a poor reputation score.

jcrawfordor 12 hours ago

There are definitely factors beyond IP being used. A week ago I found that all requests from Chrome-like browsers got a 429 across more than a half dozen networks and several machines, while Firefox reliably worked. I assume this was an overzealous policy on UAs.

BeetleB 12 hours ago

BeetleB 12 hours ago

No - my phone is not connected to work's WiFi.

Wonder who the bad actors in my company are...

flexagoon 11 hours ago

iamacyborg 11 hours ago

ButlerianJihad 11 hours ago

novok 13 hours ago

Your workplace is probably redirecting traffic through a datacenter IP range. Especially if they have their own datacenters like google, microsoft, oracle, amazon, etc.

Try making a vpn via digital ocean for example and you'll see similar patterns.

GetSMS 4 hours ago

My buddy said he could not access it even from a residential IP, it was blacklisted for some reason.

dotmanish 14 hours ago

Could be due to some scrapers from either your work ISP block, or the larger block which lends IPs to multiple workplaces.

giantrobot 10 hours ago

They seem to be aggressively blocking IPv6 source IPs. I ran into this problem over the past month traveling. I got nothing but 429 errors until I switched on my VPN (which is IPv4 only) and magically the Wayback machine worked again. The lack of transparency on the part of IA is very frustrating.

timpera 14 hours ago

I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.

Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.

emaro 13 hours ago

It's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/

I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.

zdragnar 13 hours ago

I'm a little more skeptical that this is "AI is big so it is worse" issue. Yes, AI is big in scale, but this has been the case for almost every popular free service. They either start:

- charging (news / journalist services)

- gate-keeping (X forcing log-ins)

- enshittifying (lots of ads and degraded service)

The fact that the way back machine is incredibly useful but most people didn't know about it or use it very much doesn't change the fact that it has basically become very popular... only with LLM agents rather than humans. Ads alone aren't enough to support human traffic for many sites with human traffic.

roughly 2 hours ago

Bonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.

RobotToaster 43 minutes ago

Torches and pitchforks?

CqtGLRGcukpy 14 hours ago

> We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.

delis-thumbs-7e 4 hours ago

I recently remembered a wonderful comic blog from 2010’s that is not online anymore. It was a sonderful Finnish LGTG-thened comic blog that I use to read, then forgot completely until few weeks ago. WM had it stored of course, so I could read through this amazing piece of internet art again.

I really so through some money their way, they do wonderful work.

1vuio0pswjnm7 5 hours ago

"Here's what's going on."

Thank you

https://news.ycombinator.com/item?id=49571448

I had a feeling it was due to "AI" companies and developers using "agents"

Not surprised

pelican0 13 hours ago

Is it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies?

Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.

userbinator 3 hours ago

It's not. There is no evidence, just propaganda.

sicktriple 8 hours ago

Anyone else feel like making a new internet and starting over

righthand an hour ago

People will just bring their bots over and you'll be back to square one. Bots scraping existed before LLM companies decided to go nutso on the internet.

thimabi 13 hours ago

I wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.

extralongdivisi 12 hours ago

Gatekeeping information is not the solution

hamandcheese 11 hours ago

Why not? If its the difference between the information being available at all, then I choose login any day of the week.

extralongdivisi 10 hours ago

TechSquidTV 5 hours ago

xacky 12 hours ago

The anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.

sicktriple 6 hours ago

the root of the root of all evil: NAT

HDBaseT 9 hours ago

How exactly does IPv6 "stop crawlers".

If anything, it will make it harder to block due to the vastness of the IPv6 address space.

fulafel 2 hours ago

CGNAT makes all ISP users appear to come from one v4 address, so blocking by v4 address becomes unworkable.

petterroea 7 hours ago

I'd be happy to pay a 5$/month donation to get a higher rate limit/more lenient filter put on me

throwaway456754 2 hours ago

If you get something, it isn't a donation.

ilamont 12 hours ago

Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website?

My blogs are getting slammed and there are issues with cloudflare or captchas.

iamacyborg 11 hours ago

> Shouldn't the solution be to gate bulk access for automated services for a price?

Fine in theory but determined scrapers will use residential proxies in bulk.

LastTrain 7 hours ago

These should be illegal unless users sign off on every fucking byte.

potato-peeler 5 hours ago

Wayback can’t be accessed through vpn, atleast on proton. Heck, most sites simply block you for using vpn.

xbar 4 hours ago

Thank you for the Wayback Machine. It is immensely powerful for good.

msephton 8 hours ago

Why can't they capture OS, Browser, and IP address at the time of error?

All that information is available at the point of failure, the user should not need to email it in.

ericpauley 8 hours ago

Presumedly they collect that, but the vast majority of blocks are correct and not errors. This info allows them to look up the user’s request to label it as legitimate.

tech234a 14 hours ago

I wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.

stickfigure 14 hours ago

Plenty of threads on HN about this, Anubis does not work.

danbolt 10 hours ago

I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1]

I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit?

[1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...

nubinetwork 2 hours ago

stickfigure 8 hours ago

autoexec 11 hours ago

It always seems to keep me, a normal human, locked out of any site that uses it.

phendrenad2 11 hours ago

Plenty of threads saying it works, too.

vlyan 14 hours ago

unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?

msephton 13 hours ago

They get marked as inaccessible, but still exist in IA data

vlyan 12 hours ago

is it possible to access somehow? it seems the site got excluded because of robots.txt set by some domain squatter, not manually.

msephton 11 hours ago

alightsoul 10 hours ago

tech234a 13 hours ago

See also: https://wiki.archiveteam.org/index.php/List_of_websites_excl...

Note that Archive Team is separate from the Internet Archive.

int32_64 13 hours ago

Are any AI companies using residential proxies to scrape?

xena 12 hours ago

Yes. It's impossible to tell which because the split is residential proxies, dataset curators, and AI companies all being separate actors. However I fucking guarantee you it's out there and people are too cowardly to be honest about it so they don't get sued out of existence.

oasisbob 11 hours ago

Oh yeah, absolutely.

tgtweak 11 hours ago

Can't wayback machine just offer direct access to the archive for a premium and in doing so, pay for the service?

edelbitter 11 hours ago

Not while the new dukes of the internet wielding massive armies of hijacked smart TVs have a better time browsing the web than I have; as a mere peasant with just a few IP addresses. There would be no reason to sign up and pay up for bulk access, unless open access is shut down.

MattCruikshank 14 hours ago

There was a feature on Amazon Web Services for a while, and I wish it was still there...

Downloader pays.

I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.

I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.

brador 12 hours ago

The only solution is to make visitors do compute. Compressing files for the archive to access other files would be perfect for this.

Cross verify hashes to prevent cheating.

Ez.

hubraumhugo 13 hours ago

There is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race?

Some approaches that I think are promising:

- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).

- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.

- what else?

[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...

[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...

edelbitter 10 hours ago

- Find some new way for Cloudflare to acquire paying customers. If their business did not depend on the status quo, they would be exceptionally well positioned to roll out the technical & organizational frameworks that that make massive botnets a thing of the past.

maxrev17 13 hours ago

Yeah it’s kinda crazy to me that what was once a back alley python script is now accepted as ‘fine, free for all’. The new era of bros really are smth else.

ignoramous 14 hours ago

UltraSane 14 hours ago

Why not put it in S3 with downloader pays?

Kayvanian 14 hours ago

As a public resource the hope is for Wayback to be free to access. I imagine putting up a paywall would be their last resort.

charcircuit 13 hours ago

S3 price gouges on bandwidth.

lousken 14 hours ago

AI companies should pay billions to wayback machine for access

KPGv2 14 hours ago

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

roblh 14 hours ago

Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?

autoexec 11 hours ago

They wouldn't be paying for the content, just the bandwidth. Like buying a linux OS on a CD ROM was about the cost of media not profiting off of the software.

0xDEAFBEAD 8 hours ago

Isn't that already a big part of reddit's business model?

jMyles 14 hours ago

It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.

autoexec 11 hours ago

Onavo 14 hours ago

Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.

It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.

I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

oasisbob 11 hours ago

> It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs

The problem with this perspective is that it ignores the victimization which is happening to all sorts of sites right now.

On one hand, you have content owners/suppliers which are trying to place restrictions on how much free bulk use is allowed.

When scrapers go to exotic lengths to evade the blocks, eg by using thousands of ephemeral IP addresses to collect an entire corpus, saying stuff like that makes it sound like it's all a wash.

"Oh, what a silly situation... How did we ever end up like this? It's not good for anyone ..."

No, there is a victim trying to defend themselves from rampant theft of resources, and a corporate asshole which doesn't care about the effects of their actions.

drdexebtjl 14 hours ago

Sites would just block the Internet Archive crawler as well.

imglorp 14 hours ago

Micropayments would solve so many Internet problems. It's not too late to adopt.

Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.

The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.

novok 13 hours ago

Micropayments are blocked by government money laundering regulations increasing the costs significantly to make them untenable.

Analemma_ 13 hours ago

Micropayments would solve all the problems except for the problem that people absolutely loathe micropayments. Like, vein-popping furiously hate them.

Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact that users despise micropayments and we'd rather blame credit card companies for the lack of adoption.

mindcandy 13 hours ago

imglorp 12 hours ago

xp84 14 hours ago

My guess? Because even with a paid endpoint, the type of unscrupulous yahoo that is DDOSing IA today would probably still abuse the free endpoints because they can. The revenue that might come from a paid endpoint could help to scale up, but with how slow IA usually seems, I suspect there is an upper limit to how much traffic they can serve without a LOT more revenue.

This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.

katatue 3 hours ago

IA might be large enough to earn consideration, but generally scrapers just don't care about being good citizens. I work in the GLAM space and we offer OAI-PMH interfaces for the harvesting of our collections data - which doesn't stop companies from preferring to scrape our website for worse (less complete, less structured, less standardized) data instead.

KPGv2 14 hours ago

> Why not just offer a paid endpoint for the crawlers?

Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.

Ajedi32 13 hours ago

What if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content (the IA is a nonprofit anyway), just to allow the IA to continue to serve its purpose as an archive of public data without being overwhelmed by bots.

Onavo 13 hours ago

croes 14 hours ago

It’s one thing to archive other companies content, it’s another to sell the access to it

faefox 14 hours ago

Yeah, who does the Internet Archive think it is, (insert literally any AI company here)?

bonoboTP 14 hours ago

Onavo 14 hours ago

That's for the lawyers to sort out, they have a lot of flexibility as a US nonprofit. The case law isn't that clear cut for this.

simonw 14 hours ago

celsoazevedo 14 hours ago

xp84 14 hours ago

msephton 13 hours ago

I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.

jolmg 13 hours ago

> Asking users to email them with details of their OS, browser, IP address is just crazy.

It's surely to serve as data to help tell humans apart from bots.

> Changes made by IA shouldn't become my responsibility.

They're a free service. It's ultimately not their responsibility to service you either.

msephton 11 hours ago

Imagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request.

IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.

kjs3 10 hours ago

jolmg 11 hours ago

HDBaseT 9 hours ago

doctor_radium 8 hours ago

OTOH I do appreciate their openness. It beats those times when I try visiting a site, only to get a cryptic 403 error or similar and no suggestion the site would like to hear from me.