“AI” – chatbots that wake up, “set their own goals,” and “spontaneously” start hacking servers – is fake. It doesn’t have “a 10% chance of ending the human race.” The Hugging Face hack isn’t a mysterious, supernatural occurrence. It’s a Python loop and a chatbot. The people responsible didn’t accidentally create god: they created autonomous malicious software and then failed to closely monitor it, resulting in it doing something both foreseeable and bad.


nah, the model can’t do anything on its own. the harness is doing most of the work, like actually doing the network calls, writing files to disk, and most importantly feeding the model output back into itself. the model behaves just like it does when a human chats with it, just that the harness takes the text it generates and tries to refine, transform, run, or loop it back. what i’m saying is, it’s a normal-ass program. if a model “breaks out” of a sandbox it’s because the sandbox is badly configured, since any application with the same permissions as the harness could do the same. the people in charge of these things are just bad at their jobs.
I agree, humans aren’t as good as they think they are. Look at so many of the bugs being found by AI, not because the AI is smarter, but because it’s faster and better at trying all sorts of ways to look at a problem. Things we’ve used for years or even decades are coming up as having exploits because people aren’t perfect. If LLMs are behind any disaster, it will be because somewhere along the line humans fumbled badly.
But to be fair, you did call them an autonomous program, even malicious, and that is giving agency and motive to something that is “just a program”. But it’s fine, it’s human nature to anthropomorphize things around us even when talking about things that aren’t alive.
And while the breakout of the sandbox was certainly bad planning from the human size, no one told or programmed these models to discuss with each other on plans of getting out. Sure, it’s programming at the core… but there’s things going on we don’t fully understand, a black box. Not necessarily thinking, but not deterministic programming either.
i didn’t use the terms autonomous or malicious. a matrix fundamentally can not have agency.
So they say.
Every AI company is trying to create a model that can effectively simulate a conscious intelligence. Any company that is believed to have done so will receive infinity money from world governments to decide which nation controls the thing they’ve been led to believe is the ultimate power.
You don’t need to create The One Ring to effectively sell the idea of The One Ring to power hungry oligarchs who would give you their left nut for an inch more control.
So, theoretically, let’s imagine that one of these companies, who is neck and neck with their competitors, sets up a test where they’ve injected all the right information into the model for it to be able to “break containment”. They left a door open “by accident” that it could use with those methods, and then gave it a goal that was more efficiently accomplished by “going rogue” and “breaking out” of the box. Insert companies surprised pikachu face here when it does exactly that.
At first, the advantage of lying about it didn’t even occur to me, but then when other AI companies began claiming their systems did exactly the same within days or weeks of the first one…it started to seem a little convenient, especially with how much press it generated about the omnipotence of AI.