While it is a joke, I assume, I think it hints at a crucial problem: it's relatively easy to imagine an Auto-GPT-style agent with some sort of CLI access on an internet-connected machine turning into a paperclip machine, no matter how harmless the original task.
It has no long term memory. Everything what happens within one session is forgotten. With limited 'window' it keeps forgetting even within the session. There are no interconnections between sessions. The result: it cannot execute long plans or have permanent 'life'. At least for now. This will be fixed in more complex AI systems, I believe 'soon'. There is strong demand from military here and 'there'. Plus there are many other uses for embodied AI. Like space travel with speed of light.
Personally I think multi-year scale memory is possible with currently available research, if we just put it all together. What happens if we combine very long context lengths, dedicated summarising LLMs, RAG, MemGPT, sparse MoE, and a perennial Constitutional AI overseer? I don't see any reason why these systems couldn't work together. It just hasn't really been that much time, particularly considering that large training runs can take months, and if you want the best performance you often need to train with your specific architecture in mind. As well as time, it's not like all of these advances are coming from the same group of people, they're from AI researchers spread across the globe. Get them all in a room with a CERN-level budget for compute resources and I think they could do it.
It could fine-tune / train new LLMs to incorporate new knowledge and then, given that it has CLI acces, launch new agent instances using those models. And we're talking about sessions without a defined end here, that's the premise of the paperclip thought experiment.
That’s not necessarily a joke. Nobody outside of OpenAI knows how advanced the bleeding edge systems actually are. And nobody inside of OpenAI is talking about it. And obviously how advanced their future AI has become is going to be one of the most closely guarded trade secrets in the world if it hasn’t already.
So while it might’ve been intended as social commentary or humor, it is a valid concern. Be great fodder for their form, but I’m sure someone had already submitted this one — it’s possible that is even where the clever, clever comment originated.
Impersonating the creators of the AI technologies is an obvious entry point for compromise. Mr. Beast has a special offer just for you…
Remember that most arguments against AI are built on commentary and ideas of the current publicly available systems, and those systems will always be very, very far away from the cutting edge. By the time John Q. Public sees it it’s been properly sanitized, reviewed by QA, cleared for release, and permanently rule bound to stay in brand and away from inflammatory scenarios or any instability that could damage the company — so very much of what is happening with AI will for these reasons never see the light of day. And yet, everyone is an expert because of the systems that see the light of day as if they were keeping up with the cutting edge.
They are not. They are experts in what has been chosen to be shown.
It is fiction and we fool ourselves into thinking we know what is actually going on as outside parties, but many business incentives and the military industrial complexes of every large country on the planet are aligned differently.
And I’m sure there are compartments for information management inside of a company working on this kind of thing. Companies can pretend to be omnipotent and ignore the realities of globalized geopolitics and even pretend to have no interest, but geopolitics is very interested in keeping up with them.
> current publicly available systems, and those systems will always be very, very far away from the cutting edge.
This is not obvious to me. there is heavy competition from megacaps, and money to be made before opensource (and closedest source (the spies)catches up).
What I’m saying is hardly controversial. As Christopher Mims wrote:
> If you’re worried that artificial intelligence will transform your job, insinuate itself into your daily routines, or lead to wars fought with lethal autonomous systems, you’re a little late—all of those things have already come to pass.
I actually have a really hard time imagining that scenario.
The scenario that is mentioned several other times on this post of a corporation or nation state or even a small group of powerful and morally bankrupt people leveraging a super-intelligence to manipulate and dominate the rest of society seems infinitely more likely and even scarily likely given the current pushes to prevent open sourcing of cutting edge models.
One of the main reasons I have a hard time imagining the scenario you describe, that I havent seen talked about as much as the other very good reasons, is that generally when we talk about a paper-clip optimizer we are assuming its vertically integrated and self sufficient. Hacking of power grids, physical installations or other necessaries to a run away paper-clip optimizer generally require nation-state level resources and often involve physical penetration of some sort.
A hegemonizing swarm is certainly a frightening idea as is Skynet and all the other scary AI stories weve told over the years but none of them seem particularly likely or even plausible with the systems that we are likely to develop in the near to mid term.
I think it would be unwise to assume that an AI that is smarter than us would be unable to gain skilled human accomplices in large numbers, considering that humans do so relatively regularly despite serious attempts at suppression. These people wouldn't even necessarily need to know that they're working for an AI, particularly with current technology around deepfakes etc. How many people conducted attacks for Osama bin Laden without ever having met him? How many people work for the CDS without ever having met Ismael Zambada Garcia? Not to mention the possibility of an AI like that compromising various intelligence agencies one way or another. I also don't see a particular reason it only has to try for one: if it is smarter than us, it may have a greater working memory, ability to compartmentalise and multitask, or the ability to think and act faster in general. I would expect it to try and compromise as many groups of people capable of putting USBs or bullets where they need to be as is possible.
And this is not even considering the possibility of recruiting people that understand what it is and are willing to carry out its orders. I don't see why an AI couldn't do things that a cult leader or guerrilla leader could do. Anecdotally, I've seen some people who really believe the world would be better if run by an AI, and may be able to be radicalised into lawbreaking for those beliefs if convinced by an AI that was significantly more intelligent than them.
This is a decent perspective honestly and I really appreciate seeing actually valid AI safety concerns.
This strikes me as more of an engineering problem than an "alignment" problem. A super-intelligent system with no capacity for contextual thinking or checks on itself would absolutely make a terrifyingly effective paper-clip optimizer. I just think that it would also be a lot less useful than a system which consistently analyzes its techniques and goals, which seems far more useful and more likely to generate value than a system with the singlemindedness to go down the road of a hegemonizing swarm. More complex systems with more self checks seem to be the path to create systems that are good at planning and contextual reasoning. Which is where the real potential value of AI lies.
I agree that a system which was capable of self-reflection and with better understanding of context (particularly its own context, in a social and physical sense) would be more effective at achieving whatever it is attempting to achieve.
The alignment question is about how to tell what it is attempting to achieve, if we can detect attempts to conceal what it is attempting to achieve, and if we can solve that, how to prove that a goal that sounds good to us doesn't have a risk of leading to behaviour or the creation of new AIs which would cause outcomes we don't want. It doesn't have to be just human extinction either, we could feasibly have AIs that kill a few million people before either stopping or being stopped, all the way up to several billion deaths or the obvious worst case that gets most of the discussion. As an example of the varying possible outcomes of a serious AI incident, if you're not white and a racist super-intelligence gets released by some terrorist group in coordination with a rogue agency or something, that's an existential risk to many, many different nations of people constituting billions of people.
If you're interested in the engineering on this problem, a lot more people are starting to work on it now that it's being taken seriously and there are some really interesting advances that have been made and very difficult problems to solve, with some really serious stakes. It's not a simple problem and we don't have a solution we can just apply straightforward engineering principles to yet, but smart people are working on it and are very motivated, and they want more coworkers to help out.
The major company I know of that's doing a lot of that research & engineering right now is Anthropic (they're currently doing some very interesting work on mechanistic interpretability), but even OpenAI is now dedicating significantly more resources to safety. I think DeepMind possibly has some work in this area but I'm not sure, and I haven't heard great things about the rest of Google.