- cross-posted to:
- fediverse@lemmy.world
- cross-posted to:
- fediverse@lemmy.world
Codeberg was asking about this. The linked toot by a commenter points to :
These are CC-BY-SA 4.0 remixes of the Stack Exchange Creative Commons Data Dumps. 100% Unendorsed by Stack Exchange, Inc.
They are minimal. They provide the data you probably care about and the data you need to comply with the original license in SQLite format.
robots.txt may help : https://neil-clarke.com/block-the-bots-that-feed-ai-models-by-scraping-your-website or blocking by IP addresses.
No, it can’t be. I may be using robots.txt on, say, lemmy.ml, but those posts will still be broadcasted on lemmy.world, or hexbear.net.