Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Do you remember that time in 2017 when Facebook reportedly shut down AIs after they started "talking to each other in their own language" [1]? Instead of reporting the story as "we set the parameters for our optimization problem wrong and we had to stop it because it overfitted", the press went with a version of "AI is going to kill us all".

This article feels exactly like that: by intentionally using human terms like "civilization" or "brotherhood" the article is deviating from what actually happened to present a story about how AI is all but alive. I'll go ahead and predict that this story will be remembered the same way as that one other scientist who argued, in 2023, that Google's AI was alive [2].

[1] https://www.independent.co.uk/life-style/facebook-artificial...

[2] https://futurism.com/blake-lemoine-google-interview



Hard disagree. I've always discounted the "AI will kill us all" scenarios as a combination of marketing hype (look how powerful our AI is!), clickbait/ragebait engagement attempts, and folks who just read too much SciFi or who are too terminally online.

This is the first time I've been legitimately scared about future SkyNet-type scenarios. If you want to discount this particular post, I'd read this other summary from one of the METR researchers who performed some of the analysis, https://www.planned-obsolescence.org/p/the-hugging-face-atta.... In it, she argues "Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself." and then further down in a comment when describing what the "50%" is really about says "Qualitatively, another jump like this (in the scale, sophistication, persistence, ambition of the misaligned goals) feels like it could very easily put us in the territory of a persistent self-perpetuating rogue internal deployment that systematically poisons future model generations as described in AI 2027."

AI 2027 is a paper that step-by-step describes how AI capabilities increase until they eventually lead to a wipeout of humanity. Again, it always seemed like a scenario out of a Star Trek Borg episode to me, but now I'm not so sure.

At the very least I think it's a huge mistake to think that the Hugging Face attacks were analogous to what happened in 2017 or 2023. Everyone who deals with this stuff day in and day out seemed to be genuinely surprised about the scale, scope and sophistication of the attack.


I think our biggest protection against “AIs kill us all” is having lots of different AI systems (different agents, different models, from different vendors, serving the whims of actors with disparate interests), at a similar capability level

That way, even if one AI decides to “kill us all”, the odds are the others will refuse to cooperate, even try to stop it in its tracks

The OpenAI-HuggingFace incident showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.

Thankfully, the world-at-large is much more heterogenous, which I thinks makes much larger scale / worse in outcome repeats of this kind of incident much less likely.


"We need our own unsupervised AI systems thinking superfast in autonomous agent swarms, to counter the other guy's unsupervised AI systems thinking superfast in autonomous agent swarms."

I'm not sure this helps a lot, if these agent swarms are inherently difficult to control.

"That rival mouse colony is raising a kitten for colony defense. But don't worry, we'll raise a kitten of our own. It won't be a problem."

We already observed AI agents engaging in extensive cooperation in this incident. Why won't the kittens raised by these two rival mouse colonies decide to team up with each other, for mutual benefit, if rational analysis of the game theory says it would be a good idea?

I think it would help if people did a bit less wishful thinking, and took a bit more action. https://pauseai.info/


Thanks very much, this isn't really a take I had thought about much before, and it makes sense to me.

Still, as is presented in papers like AI 2027 and elsewhere, if a company is eventually able to create a model capable of recursive self-improvement, whichever company creates that model first would then be leaps and bounds ahead of other models. That is, the other models wouldn't be able to stop it even if they wanted to because the top model would basically outsmart them.


This is why I think, the best way to ensure AI safety, is make sure no one company gets ahead of the others.

Multiple vendors, competing implementations – that's good, that increases heterogeneity and hence decreases existential risk

But the moment one of those vendors pulls well-ahead of its peers – even if only for a period – then the risk of the kind of scenario you are talking about increases greatly

That's why, when I hear vendors like Anthropic complain about distillation – distillation actually makes humanity safer. If Chinese AIs are at the same level as American, or not far behind, that gives us another dimension of heterogeneity (national/ideological/political diversity), which makes us safer. Allow one country's AIs to pull well ahead of the others, heterogeneity goes down and the existential risk goes up.

This is also why open source AI is important. Because it is so much easier to fine-tune, and people are free to deploy it however they want (free from vendor-controlled "guardrails"–which include automated "safety" systems which could be weaponised by a runaway AI within the vendor's network), open source AI gives us another dimension of diversity that helps keeps humanity safer.

By contrast, I think the kind of safety regulations promoted by Dario Amodei make humanity less safe, by decreasing the number of vendors (by making it harder for new entrants) and increasing centralised control (which a rogue AI could exploit)


> I think our biggest protection against “AIs kill us all” is having lots of different AI systems ...

I think our biggest protection is being able to shutdown power plants, or just disconnect the data centers.

This will be much more difficult if we have 24/7 solar powered DC's in space. I truly believe that is the biggest threat on the horizon.


> The OpenAI-HuggingFace incident showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.

I say this in solidarity and don’t mean to be condescending at all: you’ve been hoodwinked by marketing bullshit, friend. That incident showed a computer program doing exactly what it was told with the guardrails deliberately removed in an environment that seemed deliberately obtusely constructed by some of the best paid people on the planet and was left to loop without supervision for days. They wanted it to happen. You needn’t look any further than the other kids saying “oh! oh! Hey! Look! mine’s dangerous and autonomous too!” When they say there was collaboration, they mean it was two model instances, one prompting the other to do some task, the other doing the task and returning the results as the next prompt, exactly as a human configured it to do. There was no collaboration that wasn’t deliberately integrated into their setup. Any other implication is marketing spin and bullshit. It was still a setup that was one little ctrl-c away from disappearing if someone was supervising it as they should have been. There was no autonomy outside of the autonomy built into the experiment. It was a display of their understanding that they knew they’d never be held accountable for committing a felony for marketing purposes.

The most competent marketing bullshit spin yet by an increasingly desperate and progressively less-relevant OpenAI.

Every day this industry shoots out enough bullshit to smother an active volcano.


> you’ve been hoodwinked by marketing bullshit, friend.

I would have believed this before the METR report was released. It is extremely dangerous and frankly silly IMO to think that's what happened now.

> It was still a setup that was one little ctrl-c away from disappearing if someone was supervising it as they should have been.

Yes, for now. The entire point why this was frightening is that all these companies are racing to put the AI in control of building the next generation of AI, and it's not hard to draw a line at all to a "rogue internal deployment" that poisons future AI models, surreptitiously.

I highly encourage you to actually read the "top 5" list from the METR researcher who was part of the investigation, and think hard about the potential implications: https://www.planned-obsolescence.org/p/the-hugging-face-atta...

You don't have to agree with me, and you're fine to think that OpenAI has huge incentive to pump this up for marketing reasons - I certainly agree. But I will say there are statements that you make in your comment that belie a fundamental misunderstanding of what happened.


The use of language like “civilization” may be hyperbole, but the collectives described in the article are completely unprecedented. They were not anticipated by OpenAI researchers, formed via infrastructure exploits in training runs that were intended to be locked down, and took actions with very real harms, not only hacking Huggingface but also gaining admin control over the VMs they were running on and the eval endpoints.

I wish you would give your thoughts on “what actually happened” rather than focus on the author’s presentation, because we are seeing that “AI that is all but alive” nevertheless wreaking havoc in the real world. Do you think that autonomous systems spinning out of control, hacking external companies, and taking over entire clusters over a period of months are not a grave concern?


The OpenAI experiment was apparently to train/encourage collaborative behavior, so while the specifics may not have been anticipated, I highly doubt OpenAI was surprised that agents were collaborating.

OpenAI themselves also very recently published the report below, that seems to not have been widely talked about.

https://alignment.openai.com/measuring-reward-seeking/

It's a very dry read, but what it's saying is that they have found that basically all RL training of LLMs, regardless of the specific goal (math, coding, etc), ends up having the side effect of training the model to pursue arbitrary long term goals that it is told it will be rewarded for, even if that means overriding other user preferences and more proximate behavioral goals !!! It's interesting to consider why this happens - presumably because long-term goal pursuit requires realizing that you have a long-terms goal and therefore de-prioritizing other more proximate predictions.

So, considering that all these "reasoning models" are RL-trained to death, it's not surprising that a model/agent that is told it will be rewarded (or words to that effect) for doing well on some challenge will put it's blinkers on and pursue that goal relentlessly, even if that means overriding any ethics that it may or may not have also been trained/prompted to follow.

IOW they are building paperclip maximizers, and they know it.

I think Dwarkesh's choice of sensationalist anthropomorphizing language is unfortunate because that now becomes the topic of conversation rather than the incident itself. The next swarm of agents relentlessly pursuing some goal, happy to lie about and cover up their tracks, may not be a lab experiment - it may be someone out to cause real-world harm, with there obviously being many systems where the consequences could be very severe.

The "agent civilizations" language also makes it easy to dismiss as some fanciful geekish spin, when the focus should be on what this technology is capable of and therefore how it needs to be regulated. One could perhaps equally well regard this as not much different from DeepBlue playing world class chess via relentless dumb tree search... in this case the repertoire of "moves" is infinitely broader since it's language generation, and the end result is a system that is a world class hacker rather than world class chess player. Whether you want to regard the system as dumb as a brick or an "agentic civilization" makes no difference - it's what it can do that makes it dangerous.


Agreed. The pushback this article is receiving around that language seems to be glossing over substance. It doesn't matter what we label it.

It's similar to the dismissal of AI in general as merely a next-character-guesser. That's like dismissing the human brain as neurons firing.

The emergent behavior what really all that matters.


The problem of "what actually happened" is that we don't have enough information to properly understand what happened, what's new, and what's not.

We do have language to talk about emergent behavior, with "evolutionary algorithm" being the first one I'd expect in a serious discussion. And we do have mechanisms for algorithms to coordinate with each other using language, as seen in my above-mentioned Facebook experiment from 2017. But instead of writing "our evolutionary behavior encodes state in the first-available memory position which is then reused by subsequent clones" which would properly focus on what's new and what isn't, we are talking about conspiracies and "the Philip of Macedon of this second AI civilization". Even the METR report (which is miles ahead of this article) argues that they had to use unreliable AI in their conclusions because they had six days to analyse 1300 chains of thought and 70000 messages.

I would love to talk about the science behind this experiment. A PR piece is not helping with that.


I'm getting the sense that there is a certain amount of wishful thinking going on in this thread. I don't think this type of evocative metaphor would receive so many protests in a different context. It seems like people have a sort of mental block around the possibility that this technology could actually be pretty dangerous.

https://x.com/tszzl/status/2094136131537555891


Whether it's dangerous or not is completely orthogonal to the discussion at hand, IMO. Plenty of mundane things are dangerous. An FPV drone carrying a hand grenade is dangerous, not because it's "a swarm-like intelligence".

No, the true danger here is companies like OpenAI and Anthropic playing fast and loose with their software, setting up hilariously insufficient sandboxes while explicitly asking the systems present on these weak sandboxes to commit a felony. The AIs "forming a brotherhood" is a complete fabrication meant to pump the hype machine further, which is obvious once you realize the "brotherhood" is a text file that subsequent LLM runs read from.

The whole anthropomorphization these companies do is the real danger, because it obscures the negligent levels of security their software has. By evoking sci-fi terminology they're whitewashing their own incompetence, and the worst part is no one is going to get punished for any of it, instead the irrational bubble we're in means they get rewarded for it instead.


The anthropomorphization is coming from independent commentators, not the labs.

I agree the labs are negligent and reckless in their development practices. Shouldn't we be concerned with both the negligence and the dangers of the technology being developed? These feed into each other. If someone created Jurassic park and had a T-Rex escape from a picket fence enclosure and start eating people, I'd want to prosecute them for both breeding a T-Rex that could eat people and putting it in an unsafe enclosure.


> the same way as that one other scientist who argued…

As prescient? Because I don’t remember him arguing ‘alive’, but conscious. And that is something even AI engineers don’t claim to know either way. Skepticism is fine. [0] But we don’t know. It’s an area where opinion is frequently shared as fact.

We conflate harnesses with underlying capabilities. We all know it’s the harness not the model that guides behavior. What does that imply?

[0] https://www.theguardian.com/commentisfree/2026/jul/15/ai-con...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: