Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Others have answered your question. A lot of the sites listed here https://anubis.techaro.lol/docs/user/known-instances/ have dynamic content; Git web interfaces in particular (Codeberg, the Linux kernel, FFMPEG, and more are on the list) are vulnerable to poorly or maliciously configured scrapers.
 help



Yeah it makes sense for dynamic content. But so far I have only seen it on blogs, which could have just been a html file.

Even for git hosts, there is no reason to add bot checks to e.g. the repository root or other common URLs that random real users land on.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: