Now AI blocking will get you more precise control over your website! Now AI crawlers will be allowed to grab your website! But don’t worry! They promised and pinky swore they wouldn’t train any AI on your data. So your data is protected!

Isn’t that great!!?

  • halvar@lemy.lol
    link
    fedilink
    arrow-up
    9
    ·
    9 hours ago

    Crawlers are a big fucking problem and we really didn’t build the internet to accomodate them. Maybe if we said in 1990 “hey what if datacenters with Tbps connections will mass query all websites” we’d have a solution by now, but obviously that wasn’t really an issue back then. Anubis seems great but proof of work to prove you are a human seems to have it’s issues.

    • Natanael@infosec.pub
      link
      fedilink
      arrow-up
      3
      ·
      3 hours ago

      The solution is smarter secure mirroring for any static public content, stuff like Jekyll + Git or Atproto based sites. For dynamic content there’s no universal solution against misbehaving data centers other than blocking them

      • Axolotl@feddit.it
        link
        fedilink
        arrow-up
        1
        ·
        2 hours ago

        Wait, can you expand more sbout Jekyll + git? Seems interessing but i can’t find anything, also, why AT protocol instead of ActivityPub?