@douglasg14b

douglasg14b@lemmy.world · 1 day ago

I assume that the gitea instance itself was being hit directly, which would make sense. It has a whole rendering stack that has to reach out to a database, get data, render the actual webpage through a template…etc

It’s a massive amount of work compared to serving up static files from say Nginx or Caddy. You can stick one of these in front of your servers, and cache http responses (to some degree anyways, that depends on gitea)

Benchmarks like this show what kind of throughput you can expect on say a 4 core VM just serving up cached files: https://blog.tjll.net/reverse-proxy-hot-dog-eating-contest-caddy-vs-nginx/#10-000-clients

90-400MB/s derived from the stats here on 4 cores. Enough to saturate a 3Gb/s connection. And caching intentionally polluted sites is crazy easy since you don’t care if it’s stale or not. Put a cloudflair cache on front of it and even easier.

You could dedicate an old Ryzen CPU (Say a 2700x) box to a proxy, and another RAM heavy device for the servers, and saturate 6Gb/s with thousands and thousands of various software instances that feed polluted data.

Hell, if someone made it a deployable utility… Oof just have self hosters dedicate a VM to shitting on LLM crawlers, make it a party.

douglasg14b@lemmy.world · 2 days ago

This is assuming aggressively cached, yes.

Also “Just text files” is what every website is sans media. And you can still, EASILY get 10+ MB pages this way between HTML, CSS, JS, and JSON. Which are all text files.

A gitea repo page for example is 400-500KB transferred (1.5-2.5MB decompressed) of almost all text.

A file page is heavier, coming in around 800-1000KB (Additional JS and CSS)

If you have a repo with 150 files, and the scraper isn’t caching assets (many don’t) then you just served up 135MB of HTMl/CSS/JS alongside the actual repository assets.

douglasg14b@lemmy.world · 2 days ago

Fair fair. I missed that

douglasg14b@lemmy.world · 3 days ago

I can get a 50Gb/s residential link where I am, and have a whole rack of servers.

Sounds like a good opportunity to crowd fund thousands and thousands of common scrapeable instances that have random poisoning.

douglasg14b@lemmy.world · 3 days ago

Low key win for kink communities.

douglasg14b@lemmy.world · 3 days ago

Yeah but that was before you had billionaires of this size able to manipulate entire markets in this capacity.

douglasg14b@lemmy.world · 3 days ago

I f-ing called it:

https://lemmy.world/comment/21195810

This shit…

douglasg14b@lemmy.world · 4 days ago

This is also a strategic way to prevent resistance from Americans against an authoritarian regime.

Drones would be a significant part of that.

douglasg14b@lemmy.world · 23 days ago

It’s a play to make at home compute unachievable, forcing people to pay for subscription cloud services and cloud compute in walled gardens.