Now AI blocking will get you more precise control over your website! Now AI crawlers will be allowed to grab your website! But don’t worry! They promised and pinky swore they wouldn’t train any AI on your data. So your data is protected!

Isn’t that great!!?

  • Arola@sh.itjust.works
    link
    fedilink
    arrow-up
    14
    ·
    4 hours ago

    It’s annoying that Cloudflare is basically facilitating / enabling the proliferation of the “zero click” internet that a handful of the worlds richest people are so keen on… And while they are encouraging site owners to hold the door open for AI, Cloudflare continues to punish actual human website visitors with their “are you human” barriers to entry. It would be nice if some of the people running it actually valued the internet a little bit ffs.

    • mamg22@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      7
      ·
      3 hours ago

      Years of site load optimizations only to see it all gone by having a 20 second wait time before every other damn website. Thanks AI bros, I truly enjoy it

      I hate cloudflare’s anti-bot, it’s so slow and buggy. Sometimes it won’t load the site I wanted to, return me back to the previous one and manipulate history so now I need to press back twice or more to get out

  • halvar@lemy.lol
    link
    fedilink
    arrow-up
    7
    ·
    7 hours ago

    Crawlers are a big fucking problem and we really didn’t build the internet to accomodate them. Maybe if we said in 1990 “hey what if datacenters with Tbps connections will mass query all websites” we’d have a solution by now, but obviously that wasn’t really an issue back then. Anubis seems great but proof of work to prove you are a human seems to have it’s issues.

    • Natanael@infosec.pub
      link
      fedilink
      arrow-up
      3
      ·
      50 minutes ago

      The solution is smarter secure mirroring for any static public content, stuff like Jekyll + Git or Atproto based sites. For dynamic content there’s no universal solution against misbehaving data centers other than blocking them

      • Axolotl@feddit.it
        link
        fedilink
        arrow-up
        1
        ·
        15 minutes ago

        Wait, can you expand more sbout Jekyll + git? Seems interessing but i can’t find anything, also, why AT protocol instead of ActivityPub?

  • cannedtuna@lemmy.world
    link
    fedilink
    English
    arrow-up
    34
    ·
    10 hours ago

    Fucking hell. Why have a website anymore? Can we get a separate Internet where AI doesn’t exist?

      • Waphles@lemmy.world
        link
        fedilink
        arrow-up
        2
        ·
        54 minutes ago

        I haven’t used Usenet for anything other than file sharing, but I think i will give it a try. I kind of imagine it is like the internet in IRC format. Does anyone have suggestions for interesting newsgroups?

      • mlatu@moist.catsweat.com
        link
        fedilink
        arrow-up
        5
        ·
        7 hours ago

        needs something like treehouse from the novel otherland by tad williams. an anarchist online space, with changing entrypoints that are told only to members of said space… or to stay in reality, something like CACert, where you could get an SSL Certificate for your website for free by talking to people in person and stuff… but nowadays youd make really sure the person you’re inviting isnt some predatory submarine aibro looking for easy content…

  • MrSulu@lemmy.ml
    link
    fedilink
    English
    arrow-up
    11
    ·
    9 hours ago

    Dear Tim Berners-Lee, Did you have a backup internet that the rest of us could use?

    • corsicanguppy@lemmy.ca
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      1
      ·
      8 hours ago

      You’re gonna hate when you find HOW they generated their search data for your searching.

      • hendrik@palaver.p3x.de
        link
        fedilink
        English
        arrow-up
        6
        ·
        7 hours ago

        I liked it. I regularly have lots of niche problems and back in the day I’d just put it into Google and find someone on Reddit or whatever who already tackled it 2 years ago. Now I don’t.

        Also used to occasionally watch the web server logs and see how the search engines check for updates on the organization I’m volunteering at. Or index my employer’s website. Or the Fediverse apps I run. And tweak my visibility according to my liking. Now that changed quite dramatically as well.