Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour.

Absent that, if we assume this is just trying to bolster generic reasoning then there's no context around it that helps to forgive misaligned behaviour. If OpenAI ran these agents with safeguards off then that seems wreckless on their part. If they didn't do that, then it says the models are executing significantly misaligned behaviour even in a generic context.

Either way it seems to suggest some pretty concerning things about OpenAI's methodology.

 help



"It's okay because we did it with an Agent" is the new "it's okay because we did it with an App." Both because it's used to circumvent regulation, and because the underlying technology creates a smokescreen in dialogue among techies.

Let's imagine I made a new website but, instead of using a database, I abused some random old forum site and created new pages on that forum for each row of data. You'd call that abusive, yes? I'd be an asshole, yes? And the fact that my website was really cool and techy would have no sway on the fact that I'd be an asshole, yes?

Well then why does OpenAI's abusive behavior get discussed in these terms? Whether it was a "reasoning type task" or whether they "instructed misaligned behavior" is irrelevant. Nobody should care. Discussing OpenAI's behavior in these terms is just a distraction from the problem at hand.


If you actually accomplish something like this your post will be on top of HN and discussed with reverence.

Source: Every Tom7 video.

I am not saying someone trying to run this as production would not be an asshole, but the technical feat is amazing. I don't see the difference between Tom7's harder hard disk video and this. Of course this is a bug in the agent but it's a fascinating bug and no one is being an asshole on purpose.

Now I will wash my fingers with bleach because I just defended the OpenAI.


A harder drive made out of neglected wikis and forums. I hate it so much I might actually try to make it, just to prove a point.

Interesting. So there’s no “they were told to hack” excuse here.

There is something fundamentally wrong with their reward function, this is pretty classic paperclip territory. And even knowing that, I expect we’ll need to see legal action with teeth against the labs before changes start being made internally.


From the report, they also tried to impersonate the moderators and perform XSS attacks (report says "unclear why they would do this at all"). So not just using a static message board either, but actively interfering with oversight.

OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt.

Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.


It's almost as if it's not actually possible to align an unknowable mystery box of floats.

Good thing we're not trying to deploy them into fully autonomous weapons or anything....

I wouldn’t take the fatalistic stance that it’s fully impossible - but it’s certainly impossible to align a model while racing as fast as any technological paradigm shift has ever raced.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: