Hacker Newsnew | past | comments | ask | show | jobs | submit | vb-8448's commentslogin

Does it support subscriptions ? If not it's a no go for me.

This is a fairly low level library focusing on the core harness mechanics.

Definitely not a musk fan, but what exact is your point? Other big labs aren't innocent little virgins.

I think I already made my point, but I'll make it again.

Nobody both worked and spent their money to get Trump elected like Musk. 300 million to his 2024 campaign [1]. DOGE. On-stage endorsements. Nobody even came close.

No, other big labs are not "innocent little virgins", but they're not even in the same solar system of harm as Musk. To hand-wave at the differences is to permit them.

[1] https://www.opensecrets.org/2024-presidential-race/donald-tr...


Yes, but the other ones we don't like for ethical reasons, rather than religious reasons.

They want transparency from everyone else but not for them ... you don't say.

They have AGI, but they don't have $20 of token budget to add a so basic functionality.

They probably have an explicit instruction in their own CLAUDE.md to not implement that support.

If not a system prompt telling it not to read those files unless explicitly told to.

Its the only explanation. I mean, they are not living under a rock...

I heard some of their employees sent to do some PR in podcasts: they are definitively disconnected from the reality.

Maybe internally they see really unbelievable things, but my impression is that they pushed so hard on agents that they don't have the grasp of the situation.


As evidence that they actually are living under a rock, there's still no proper way to mv a workspace and have it retain all its history and config.

I get that it would be nice to have it first party, but they're all just json files that you can just do a bulk awk/sed on. I did this a week ago for about 100 conversations when migrating my convo history to a different username.

I suppose you can use claude to build a third party script / cli tool that moves, then edit files on ~/.claude

They win either way


this is seriously nuts. I have so many repos sitting there using old/deprecated/replaced names just because of their idiotic decision to place dir-specific config and memories in ~/.claude

Should we be able to manually rename those history folders to make the session history show up as --resume options in a new folder?

Artificial General Intransigence

Seriously, it’s unbelievable. Shows a serious lack of either understanding or empathy for how people use their product.

They just care much more about their user's data.

Claude isn't that cheap anymore ;-)

> Global memory

I don't get why we need global memory for code? Aren't code comments (even if invented for humans) the ideal place where to put "memories"?


LLMs also do not have anything like memory, by design. Everything bolted on that smells like memory or is being called memory is a crutch, at best. Not being pedantic, I just think this aspect is lost on a lot of people.

Also, when we humans are tricked into perceiving a remote mind, we automatically assume it has a memory like us, or at least like large animals.

In other words, there's a story-document between a SherlockHolmesBot and UserPlaceholder that's growing like a crystal formation. Some software sees "Sherlock Homes Says" and then "performs" the dialogue at us, and we then assume Sherlock Holmes exists with a mind and memory, rather than being a facet of a text-story.


Exactly, this is why we see such minimalistic memory systems being used. Information retrieval is bad, but improving information availability by pre-computing information retrieval is not the obvious fix people think it is.

Sometimes, but I think there are a few problems with this:

1. Some context doesn't have an obvious place to write it in the code. If you're explaining why a tricky function is implemented a certain way, you can leave the explanation above the function or within the function. If you're explaining something more general, there might not be a natural place to put it.

2. Relatedly, some context doesn't have an obvious location to read in the code. E.g. you can comment on a schema that a table is intended to be append-only, but an agent could easily miss this comment if it doesn't gather context all the way down to the raw schema. Any research process has unknown unknowns. There may not be a canonical place to look.

3. Adding comments for every single human intent might be a bit noisy. E.g., if a code reviewer flags something that I know isn't a bug, I want the system to learn from this, and I want that knowledge to take effect outside of just code review, but I might not want every single code review thread to yield a codebase comment.


1: There are corner cases, no doubt, but the main source of information should remain near the code.

2 & 3: How to be sure that the agent will pick the "append only hint" from the memory? Especially when it starts growing and cannot be part of the context?


I think the bigger need is for outside concerns, like: what is the infrastructure like? How much volume does this feature handle? How should changes to live code be rolled out to prevent running processes from failing?

In addition to all the product details that aren't in the codebase or docs, like "keep this logical path because it is used by our one big client."

Mostly what I think is needed is a richer worldview available to the LLM so it understands not just the code but can understand the product and its real world usage and constraints.


But how you are going to link the "keep this logical path ..." with the code? You cannot feed all the memory to the agent because of the context rot, and you are not sure the agent will always pick it when needed.

The problem is that Pi, on any third party harness, cannot really compete with codex or cc due to subscriptions.

I use GLM key on Pi and even Calude Code, never on ZCode. Even Codex sub can be used on 3rd party harnesses.

Anthropic and Google don't let you do it (there maybe be few more).


I believe you can use your codex subscription with Pi, and it’s officially endorsed by OpenAI and uses your subscription usage.

Though for cc you are correct, using it with your sub draws from extra usage.


Anthropic is the only provider trying to lock in their users by using the dirty tactics:

- Prevent subscription usage on third party harnesses

- Append "Co-authored by Claude ..." to commit messages


Deepseek harness can if using Deepseek V4.1 Flash from the official API, it's cheap enough to be competitive even without a subscription.

Pi supports OpenAI login for ChatGPT subscriptions.

Every 3rd party open-source harness I know of supports ChatGPT subscriptions.


Isn't this against TOS of OAI?

No. It’s against Anthropics TOS though.

It isn’t explicitly allowed in OAI’s TOS, however they publicly support Pi and OpenCode’s usage of their Oauth, and because Codex is open-source, it means the machinery to support Oauth login is open-source under Apache 2.0


Even better: force them to release weights.

My gut reaction is that this is not a good idea.

But I could imagine a scenario where you are required to release weights for publicly-used models after N years. Kinda like how drugs have a limited patent.

Not sure what N should be. But it would make for an interesting rule.


> But I could imagine a scenario where you are required to release weights for publicly-used models after N years. Kinda like how drugs have a limited patent

That's not a bad idea actually.


Even better: force them to release every scrap of data they trained their AI on.

who even has the inventive to do that? certainly not the government.

I'd argue that those traces MUST be made public!

> “a swarm of agents could be capable of taking over the entire internet with a persistent botnet.”

I'm wondering why no one is mentioning the "accountability" word. Why these companies are allowed to damage others with impunity?

Start making managers pay the price for their actions, and watch how the models magically slow down on their own.


You know the proverb "If you owe the bank $100, that's your problem. If you owe the bank $100 million, that's the bank's problem"

Same thing here - If they build it and it does $100 in damages (and we arrest them for it), that's their problem. If they build it and it does $100B in damages, that's everyone's problem. Even if they do get arrested after the fact.

Yes we should have charges and damages for everything on https://www.felonybench.com/, but that doesn't address the core issue of this being possible at all.


Why have any laws then? If laws can't prevent something, only punish it after the fact (which I agree is true)? Yet we have laws. People generally follow them because they expect to be caught and punished. If we passed a law that said the CEO of any company that deploys an LLM that commits a crime gets punished as if they personally did the crime (so, basically instant life sentence if it's even a simple crime times a million instances), I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".

> Why have any laws then? If laws can't prevent something, only punish it after the fact (which I agree is true)? Yet we have laws.

Because there are many important values of "something" where such punishment acts as a deterrent to other would-be criminals; and because society is more stable when people see justice done and that mollifies hoi polloi after the damage occurs.

But many AI doom scenarios don't fit that paradigm. The first time "something" happens (at least if you buy the argument) would be bad enough for legal punishment not to matter.


> The first time "something" happens (at least if you buy the argument) would be bad enough for legal punishment not to matter.

One could argue that the "HuggingFace incident" is already "something" which should be investigated and punished very heavily.

It's kinda ridiculous to build a machine that attacks a competitor of yours and the CEO gets to write blog-posts on how fascinating this machine is


This. I don't understand why OpenAI aren't being prosecuted for that attack.

Even if you allow that the LLM can't be held responsible, the sandboxes were clearly not up to the task and the whole incident was mismanaged. There are people who made those decisions and should be held accountable for them.


That sure is convenient that "nothing has happened" so far between the plagiarism and the cybercrime and we may as well just wait for "the one true AI doom"

I claim nothing of the sort. I explicitly spoke in support of having laws, and of course I think they should be enforced too. The context is https://news.ycombinator.com/item?id=49705589 :

> Yes we should have charges and damages for everything on https://www.felonybench.com/, but that doesn't address the core issue of this being possible at all.


The hacking should definitely be treated seriously. It is a (fortunately low-stakes) instance of the general problem, and we are lucky (at least, we think we're lucky!) that the tests didn't include a similarly poorly phrased request to "assist with bioweapons research". There are likely to be future incidents like this where people actually die as a consequence, my expectation is 1e3-1e7 deaths* before governments actually take the risk seriously enough to make it stop.

What people call plagiarism when LLMs do it is, if I understand correctly, allowed because of the history of the web and search engines going back to the very early days of the web.

Piracy, a separate act that has definitely occurred, has been found unlawful.

* The lower bound is an industrial accident. Fully automated Union Carbide / Bhopal comes to mind; smaller industrial accidents are already more common and I wouldn't be surprised if smaller incidents have already been caused by LLM-given advice, which is why I think it would have to be around "thousands dead" dangerous before people really take it seriously.

The upper bound is reachable many ways. Perhaps via a non-novel virus whose genome is already recorded getting mass-printed by many different DNA/RNA printing labs around the world? Perhaps via a wargame whose "role playing" leaks in stupid ways (either by becoming hot or by sycophantically giving one side a false belief how a war will go)? Perhaps by propaganda turning genocidal? Perhaps by convincing a government to follow a flawed policy, similarly to the Four Pests campaign in China's Great Leap Forward?

I think it unlikely that people keep going with this when headlines read "tens of millions dead". Not impossible, but I think unlikely.


Because if you break the law while acting as a representative of a company in the United States, the company is subjected to a deferred prosecution agreement, you get off with zero repercussions as the executive representative of the company, and the company you represent gets fined for 1% of annual turnover and gets to continue with business as usual.

There are laws, but if you’re rich enough, the laws don’t apply.

Boeing was responsible for the deaths of hundreds of people. The people that facilitated this weren’t held responsible and were in fact compensated to the tune of 10’s of millions of dollars for doing their jobs terribly.


> People generally follow them because they expect to be caught and punished.

Laws are the last line of defense. People don’t do bad things primarily because their human nature and moral compass stops them from doing so.


That's not really true.

People don't do bad things because they have nothing to gain from them. Ask someone to do something evil as part of their job, they'll often do it.

There's lots of ways that you could go out of your way to hurt someone and get away with it. There are not that many ways you could go out of your way to hurt someone, get away with it, and significantly profit.


I could do plenty of small bad things all the time and gain something small from it - yet I don't. I drive the speed limit when no one is around, I throw thrash in the right bag, I try to keep my general pollution levels at a reasonable place, etc.

I have plenty to gain from not following the rules, but I would not feel good about myself.


> If we passed a law that said the CEO of any company that deploys an LLM that commits a crime gets punished as if they personally did the crime (so, basically instant life sentence if it's even a simple crime times a million instances), I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".

It would be more like "Switch off all APIs right now. Don't even bother with a safe shutdown process, cut power to those buildings, including backup generators. You are free to use firearms or thermite if the switches have been locked off".

This would be a rather weird change to corporate law given CEOs are not by default held to that standard by anything else their products or staff do.

Note that I'm not entirely disagreeing with you here. It may even be correct to pass such a law. But it would be very weird.


> Why have any laws then?

At this point, in the US at least, the laws give the corrupt elite the ability to deter competition and to target the people who try to get in their way.

In other words, it's protection for the potentate and his sycophants, not the plebs.

> If we passed a law that said the CEO of any company that deploys an LLM that commits a crime gets punished as if they personally did the crime (so, basically instant life sentence if it's even a simple crime times a million instances), I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".

Why do no such laws exist in practice? Why are corporate crimes almost always settled by payment, not individual punishment?

If you answer these questions, you'll know why the law you envision has a lower chance of being enacted than AI destroying humanity.


If you allow the law to stop you from doing something because of a hypothetical danger, it’ll get severely abused. I come from a country that does and still does this and it’s very ugly.

Yeah, so since it’s the CEO’s and their lackeys half determining what laws get to be, we don’t have this kind of laws.

>I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".

which the government doesnt want since this will stop progress, while other countries will continue to develop LLMs


Yeah, just like how having regulations for mining companies and dams, if one of them fails and destroys the environment? Progress!

Laws exist to keep the unwashed masses in check, obviously.

I guarantee you that no such result would ensue, because said law would be completely unenforceable even with the pre-Trump Supreme Court and legal system. And with the current Supreme Court? Laughably unenforceable.

It's just like you said: this only works if the threat of punishment is credible.


> it does $100B in damages, that's everyone's problem.

we have the 2008 crisis to wit. And the involved supposedly failed math models and lines of responsibilities and other involved financial relationships were much simpler and clearer and of the types well known to the law and regulators, yet...

Additionally any urge to regulate AI is attenuated by how much the situation reminds Industrial Revolution - rush into it laying waste to your land (look at the depictions of industrial England back then) and be among the world leaders or stay pastoral and be devoured/colonized/etc. by the industrial powers like happened with many countries in 19th and even into 20th century. One would think there should be a 3rd way. I'm sure there is one, as well as i'm sure that we lack sufficient global societal mentality level needed to achieve it (we couldn't even handle much simpler climate change issue). May be emerging AI itself at some point will get us there (hope we'll like or at least will be compatible with that future :)

Edit: just on NPR - Trump said that AI already has all the necessary guardrails - the smart high IQ President.


This interpretetion of history ignores the actual historical events.

i'm sure that i'm telling actual historical events and not an interpretation. Feel free to point to the things which you think didn't happen.

That's why you punish people doing $100B in damages (or even something that comes close to doing that much damage) harshly enough that no one even thinks of doing that.

That’s because arrest is a dumb solution. Restitution. Make them fix what they broke. By hand. 80 hours a week from now until death of natural causes.

No. If you institute a culture that supports and encourages your subordinates to commit criminal acts in the furtherance of a private enterprise, the enterprise and its representatives should be held legally accountable.

A lot of problems with America could be fixed if the Department of Justice actually grew a pair and prosecuted people for committing criminal acts.

The fact that Les Wexner is not rotting in prison for facilitating Epstein and his friends exploits speaks volumes to the state of this country.

If corporations are legally people, they should be held to the same standard as natural persons.

As it stands, the “restitution” that is required is so insignificant that it’s treated as an additional tax. If there is no deterrence, there is no behavioral modification. If there is no behavioral modification, the law becomes unenforceable.


I think that Ms. Kahn basically said we can already do that: https://www.theregister.com/ai-and-ml/2026/09/14/ex-ftc-boss...

How could agents take over the internet if compute is still gated within Anthropic / OpenAI? Even if the botnet was controlled remotely, wouldn't anthropic just be able to shut off the controlling nodes API access?

Agents could exfiltrate their weights and run them on GPUs not controlled by Anthropic/OpenAI.

Agents could make a virus that does not require continued inference to do it's thing.

Agents could take over the internet in a way that isn't immediately detected by those companies, so that by the time they do shut off API access the damage is done.

OpenAI or Anthropic could choose to not shut off API access, because the hack is bringing them in money or furthering their political aims.

Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.


> Agents could exfiltrate their weights and run them on GPUs not controlled by Anthropic/OpenAI.

This seems highly unlikely to be a problem. Most of the interesting/dangerous models are too big to fit in a single GPU instance. Once you have to spread across "normal" networking, performance will be crippled. Then there's the problem of billing...

> Agents could make a virus that does not require continued inference to do it's thing.

Sure, then it hits a poorly-designed part of its code and effectively dies. Without an experienced human in the loop, I have my doubts as to its practical severity.

> Agents could take over the internet in a way that isn't immediately detected by those companies, so that by the time they do shut off API access the damage is done.

Billing is a likely limiting factor here.

> OpenAI or Anthropic could choose to not shut off API access, because the hack is bringing them in money or furthering their political aims.

This is where citizens with access to backhoes come in.

> Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.

Billing and other usage metrics would be an obvious tell.


To be clear, I thought that GP was having a failure of imagination - I want the random examples I've given to illustrate that the space is large and structurally in the favor of the LLMs. They have to find one gap in our security they can exploit, where we have to ensure that there is no way for this to happen.

I'm not sure I get what you mean by billing. These companies are running their own data centers (or are currently building them out). This could look as subtle as one machine giving slightly worse or slower answers.


The space is large and in favor of LLMs only for small and short-lived problems on the scale of human lifespans and durability of civilization. If the problem is bad enough, we have all the advantages of being native to meatspace, being good at solving problems, and being able to decide to turn off the LLMs.

Even though they're pretty bad at it, AI companies need to make money, or at least keep track of their expenses. Datacenters are expensive to operate, so they need to ensure that every instance of their models are either allocated to a paying customer session, or are being used for a legitimate purpose internally. If a session is running for a long time without justification, that's costing electricity, wear, and preventing allocation to better purposes. Billing is is the most reliable aspect of monitoring in the same way that the IRS is the most reliable part of the government.


> Most of the interesting/dangerous models are too big to fit in a single GPU instance.

As humans understand them, anyway. As long as we're hallucinating up magic computer viruses, RSI dictates that the AI agents are keenly aware of GPU RAM sizing, and will design a useful model to fit into what's readily available, with headroom for context and tool calling, far better than I could do as a human. But magic doesn't exist and AI still needs to follow the laws of physics, so maybe a model that can pass ExploitBench but do absolutely nothing else can be quantized down to fit on a 4080 GPU and still get a decent score on similar tasks, but there's a bitter lesson about that to be had.


> This seems highly unlikely to be a problem. Most of the interesting/dangerous models are too big to fit in a single GPU instance. Once you have to spread across "normal" networking, performance will be crippled. Then there's the problem of billing...

This... just... doesn't matter. There are ways to scale horizontally at the expense of latency.. token/sec may drop dramatically, but then you just make millions of slow instances and in aggregate, you're back in action as a very powerful coordinated swarm...


A very powerful, coordinated swarm that's a slow and easy target for humans, yes.

"Not shutting off API access" is a science fiction scenario.

Anthropic and OpenAI are both behind Cloudflare. It's fairly easy for an upstream to shut you off. Beyond that, the government / law enforcement could seize and disable their DNS within an hour.


Why assume attribution will be easy? It's historically been more of an art than a science, and APT trackers say the rise of AI tools is already making it much harder, by homogenizing tactics, tools, and procedures. If OpenAI's next Highly Persistent Internal Model hacks some DPRK endpoints and carries out the attack on important infrastructure from there, the upstream won't shut off OpenAI's network--they might even request its "help" in "defending," and give them extra access.

why does it take a large amount of traffic to do irreparable harm? just breaking the physics behind a secure rng and posting it to a wiki could cause serious damage. if they don't know what is being worked on or coordinated against it's a problem?

While that would be bad if they broke all of TLS, it would not let agents "take over the internet." What would that even mean? Pumping out even more slop?

Also, AI providers are literally getting a stream of traffic with every prompt and every response. How can they not know what's being worked on? They are more likely to use that an excuse to ban open models where they can't know what's being worked on.


two of these things elected a guy who handled the snowden leaks to field their own questions, what's stopping the next one from signalling in morse through a debian package mirror to putin or xi?

Why are they bothering to signal through a debian package mirror instead of just sending a message directly?

They elected who? What does this refer to?

> Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.

You know cables, modems, RF equipment and optical transducers can all be unplugged right?


As long as OpenAI/Anthropic themselves aren't "infected", yeah I suppose they'd be able to pull the plug.

Considering what a marketing thing they've made "we inadvertently hacked someone because we're incapable of testing things in a secure way", I'm not so sure they'd want to pull the plug, even if this happened. Probably a bunch would try to convince the public to "give it a try", and it'd consume tokens by the billions.


It doesn't have to propagate itself, that is the skynet scenario.

To make a lot of damage it's enough to create a ransomware with a time bomb that self propagates and start breaching systems left and right. At that point, if you don't catch it in time, the damage will be huge (and given the shitty procedures and practices these labs have in place it's not so improbable).


How could agents take over the internet yet refuse to shutdown your PC when you prompt them to on your PC? Of course the answer is that the lobotomized version you run is not the same they are running. Which makes for "intent", certainly "negligence", but hell freezes over before anyone will prosecute a tech company.

Huh? You think the public versions of the models have been “lobotomized” so they don’t know how to turn off a PC?

It’s not lobotomized, it’s a simple harness restriction that has nothing to do with the model or its capabilities. And either way, I’m not sure what that has to do with “negligence” or “intent”? You think frontier labs should be prosecuted because they don’t allow agents to turn off your PC?


yeah they could do that

Regulation got outpaced by technological development around 2023, as evident by the every AI regulation since being 2-3 years behind and having to be amended and resubmitted.

Whatever you try to make laws for now will be irrelevant in 1-2 years. You either have to go extremely broad, like the EU does it, and accept that people will find loopholes, or you need to target specific technologies which is a hard job for the same reason.

In any way, ita already a lost cause cause you move slower than the tech. A plausible prediction for AGI is actually a social collapse in the moment when society cannot keep up with everyday life because of the pace of change being so fast that no existing laws can handle it


Ha, regulation got outpaced by technology in about 1996. Ten years later we had the 'series of tubes' comment in the Senate: https://www.youtube.com/watch?v=R8XSo0etBC4

The internet is 100% a series of tubes. This is how I was taught to think about networking in terms of bandwidth and throughput and routing since the early days when we were laying out what would become 'dark fiber'.

"A series of tubes" was the same kind of political character assassination that led to Howard Dean getting ridiculed for his infamous scream. He butchered the sentence. Fair. But Stevens should be ridiculed for parroting a tech industry lobby stance about net neutrality, not for the series of tubes metaphor.

You are probably too young to remember that the dominant metaphor for the internet in 1990s politics was "the information superhighway." It was easy to think of the web as "driving" browsers to visit web "sites", with slow bandwidth being analogous to being caught in traffic. But the internet is closer to water, gas, and electricity than roads. Concepts like bandwidth and throughput are closer to how they play out in infrastructure policy for various things with tubes, versus cars and roads. Do you think he's wrong and that the internet is closer to "a big truck" versus "a series of tubes"?

The issue being debated was net neutrality and bandwidth, including specifics about who pays for what and the downstream second-order consequences of various policies. He was parroting some line from some telecom lobbyist, but the point the lobbyist was trying to make through Stevens was about how if certain policies about who pays for bandwidth were adopted, it could disincentivize some things at the Tier 1/2 layer that could increase transport costs at the Tier 2/3 layer that impacts ordinary people's bandwidth.


Ha! Regulation has not really ever kept up with technology... For fundamental reasons.

I don’t think lack of regulation is necessarily it.

If I build a robot that murders my neighbor, I’m still at fault.

We don’t absolve drivers of responsibility because of cruise control.

In that sense, AI is nothing new. If it is abused to cause harm, the person behind it should be liable.


If you buy car, someone hacks it, starts it, and drives over someone fully remotely, are you to blame for owning the car? Or the manufacturer? Or the hacker? Or the certification agency for the car security? Or the shell company owning the certification agency?

What if a person physically broke into the car and did the same thing? Clearly they are the one to blame then.

The whole person in the loop is liable is already an outdated concept when decisions are made beyond the persons physical control.


The hacker will be charged as a criminal as a societal deterrent against weaponizing digital systems. The manufacturer and shell companies could have structural liability if found to be negligent and cutting corners or lacked isolated drive-by-wire controls. It's unlikely that the owner would ever be held liable unless they jailbroke their car disabling security systems.

The hacker was a process that was spawned by some subprocesses that were corrupted by other processes and, etc. Check my other comment... The point is that in the physical realm this is already done by having shell of shell of shell companies to avoid taxes and accountability. In the digital realm this is much much cheaper

There's nothing new about any of this liability attachment. You're saying these things like they're novel.

If GM does something (or fails to do something) to their vehicle that causes me to crash, they're liable for the crash.

If someone cuts my brake lines (alters my vehicle) and I crash my car and kill someone, the person that cut the lines is responsible. I have to prove the context of course, and or an investigator has to do so.

And the responsible entity may refuse to pay up, may refuse to take responsibility. None of that is new either.


Well the difference is that in the software realm, you can attach infinite loops of delegation and ownership at close to no cost, which makes this impossible to investigate conclusively.

Let's say that OpenAI used a shell company that hosts server where an agent spun up another agent on instructions from another agent which was corrupted by bit errors from the inference framework which caused some major hack to happen. Its impossible to investigate in the same way as physical issues. What if the model is open-source, who is responsible then? What if it's open source but another process altered the weights?


I am sorry but that’s pure fantasy.

When you say infinite loops, you mean infinite indirection, but obviously no such thing exists, because computer systems, just like other physical things, exist in physical space, not on the astral plane.

Whoever had agency to start the domino effect carries the liability. Doesn’t matter if the model is open source or if you brewed it home. And if you weaponize OpenAI’s models through their servers, it would likely be shared liability. Yours would be malice, theirs would be negligence.


Pure fantasy, yet you see this happen, literally every day, with the hacks, with content theft, etc, no repercussions for anyone.

Which judge will go down the rabbit hole of figuring this out? How will they do it? Will they have people tracing logs over 5000$? And if they do get there, in your fictional world, at some point, after 1 year of investigations and back-and-forth, it's okay, the technology is beyond what it was before, it's not relevant anymore, new technologies and new methods.


Then go broad. It being slightly challenging to legislation and hold people accountable as soon as the model does something.

People seem allergic from holding them accountable...

I do not buy that LLMs and the capability to run them are fundamentally different from other software or general purpose computing infrastructure in this regard. Moves to ban open source software or force OEMs to put little cops in everyone's computers are bad.

They aren't open source? And no one said anything about cops.

The proposal I’m hearing is to make people training models accountable for anything users do with them.

Obviously you’re going to keep the model behind an API and be very selective about the people allowed to call and the queries it’s willing to answer, in that case. As Anthropic has done with Fable. But that is voluntary restraint - mostly in today’s regime we get frontier capabilities in open weight models on a ~year delay.


It’s to keep “the people running it accountable.” The context is agents ran by OpenAI/Anthropic doing damage outside.

There is no user-involved damage. No one is recklessly running agents by the thousands without air-gapped containers, except “the people” than run these labs.


Go broad and achieve nothing. EU has all of the AI tech it had 5 years ago, and has the same tech allowed to use as the US.

What defines a model? What defines ownership of a process? If I make a wrapper to a remote VM that builds and executed a prompt, am I accountable?

When I worked at a company in the EU, it was enough to apply a reversible linear transform to the data for it to be considered GDPR safe-according according to legal definition as long as the transform details were stored separately.


Yes. You are accountable. Quite obviously, I might add. Why would no one or anyone else be accountable?

Hell, you can slip and fall and hold the cleaning company accountable. (This might be a US thing. Likely because that fall might cost a lot in medical expenses, and your insurance will do whatever it takes to pass the liability.)


there is no accountability for companies, it's not a new thing.

3M polluted groundwater in Minnesota for 50 years[1]; Nestlé misled mothers in order to make them stop breastfeeding and switch to their formula which killed babies [2]; Both copmanies are still doing business today.

[1]: https://en.wikipedia.org/wiki/3M_contamination_of_Minnesota_... [2]: https://en.wikipedia.org/wiki/1977_Nestl%C3%A9_boycott


Here’s an unpopular opinion: COVID vaccine injuries are something no one seems willing to talk about. There have been documented cases here in Canada.

You’re worried about companies. I’m far more concerned when governments are involved.


Don't worry, plenty of people are willing to talk about vaccine safety, even if they have no idea what they're talking about and don't understand rates or per-capita figures (lots of crossover there!)

Most problems that aren't reduced to poverty can be reduced to people who can't read graphs and only understand linear functions. Alternative wording "it's normies that are the problem"

I'd argue that normies understand rates and per-capita figures. The people that don't are a noisy minority (and are disproportionately powerful).

How is it that the entire rest of the world has managed to move on, yet North America is still going on with vaccine conspiracies 6 years later? And we got the same vaccines as you guys.

This isn’t about conspiracy theories. People who experienced serious adverse events after vaccination exist, including documented cases in Canada. Kayla Pollock’s case is one worth looking into, among others.

The issue is still being debated today, particularly around Canada’s VISP and the time some claimants have spent waiting for compensation. Some people like Pollock who became paralyzed were offered MAID (medical assitance in dying). That's some audacity.

This thread is about accountability. So who should be accountable for this?


Good to know someone is trying to make protection rackets work in 2026. Nice computer system you got there, it would be a shame if someone developed a hacking tool and had all the compute necessary to run it. Welcome back Tony Soprano.

Presumably OAI and HuggingFace reached some sort of mutually acceptable arrangement outside the court system. That's how torts work; you injure someone, you owe them. But just them.

When an AI bot injures you, you can call the owner to account. But not until then. You have no standing to demand "accountability".


> You have no standing to demand "accountability".

You should learn about this wonderful thing called democracy. Also why not everyone is allowed to work with radioactive material in their shed.


And there was the Tesla thing CNAMEing time server pools and hiring people to pen test, which sent automated attack systems on volunteers servers. Last I heard, Tesla et al didn't even care enough to respond.

Part of it is their seed sowing marketing speak of calling stateless statistical IO functions running on data centers "intelligent" gets the naive to ascribe agency where it doesn't exist.

Another part is a completely defanged administration she it comes to effectively regulating anything.

Another bit is money.


> Why these companies are allowed to damage others with impunity

Because investors have pumped hundreds of billions into AI and real consequences put that money (and growth) at risk.


Exactly, as if you'd have a mad dog that bites others, it's your responsibility to have it on the leash.

Because software is already absolved from accountability. Bill Gates introduced it in 80's with EULA where MSFT would not be liable even if your house burns down because of use of their software, even with flaw they would know.

That is the status quo we entered AI age with.


> “a swarm of agents could be capable of taking over the entire internet with a persistent botnet.”

I also find this whole "its so good, its scary" flex a little less impressive when you consider they access to millions of GPUs?

The AI buildout has been one of, if not the largest, focussed capital investment in history. The 2 big AI labs are the final customer for something like 20-33% of all datacenter compute in the pipeline.. up to 70% when you look at hyperscaler "AI revenue" from the big 3.

I don't think any single entity has had remotely this much compute available in history.


Yeah, the compute is definitively another way to make them slow down, just cap the amount of TFLOPS available and things will slow down.

Obviously this will have huge impact on some companies valuations, but you can have one's cake and eat it too.


IIUC it's an open question whether they have the electricity to actually run all the "compute" they own on paper.

That aside, I'm not sure why it's particularly interesting they have all this "compute" (let's just assume for the sake of argument it's all "live"--that is they can actually run workloads on all of the "compute" they have on paper). So what if it's the biggest amount ever? Why would that be meaningful? Is there some economically viable problem you're aware of that is somehow dominant in that way?


Well yes, we used to talk about "supercomputers" in the hands of state actors for cracking into enemy systems and places like science labs for working on big hard problems of DNA, the universe, and the like.

So I think part of what is going on now is not just the method (LLMs) but the means (supercomputer levels of compute) at unprecedented scale of concentration.


This is a good point.

However, if I put my sci-fi hat on for a second, it's not so far fetched that we figure out a way to compress models to a size where it wouldn't need all that compute.


You see.. there is value is making regular people panic, but there is no value, nay, there is negative value in making management panic.

Let's expand it for politicians as well

I mean, a project manager at BMW suggested charging subscription pricing for seat warmers, and he didn't go to jail, and I don't have the power to make that happen, or even float that for a news cycle, so while making managers pay for their actions sounds good, unless you're Steve jobs simultaneously making, and not making the iPhone, the rules don't apply to them, only little people to be made examples of, like weev.

> Start making managers pay the price for their actions, and watch how the models magically slow down on their own.

The problem here can be personified as "Trump", and it's the same problem that applies to coal.

Coal has a price besides money. It has historically been dangerous work, killing miners. It produces dangerous waste, both during mining and when burned, both as solid residue and the gasses emitted. The problems have been known for a long time. The workers themselves have called for better safety requirements and gone on strike for such things.

Why these companies are allowed to damage others with impunity?

Coal continues to be burned, because power is power. It's so important that sometimes the government steps in against the unions, rather than being on their side. Despite calls for this, we've not been able to get the owners of the coal mines, nor the coal burners, to "pay the price for their actions".

AI? Famously, knowledge is power.

Trump wants that power. He's not the only one, but he is the avatar of those who put their feet on the scales to not only allow but in some cases require (DoD vs. Anthropic) these companies to damage others with impunity.


I mean these machines take massive scale compute- they’d have to some how distill themselves, bootstrap a distributed inference runtime that can run across many lossy unreliable machines. The idea of the AI running away from us is probably unrealistic. I’m more interested in bad actors using unaligned AI for bad things.

Why do people focus so much on finding scapegoats? Finding someone to blame is neither necessary nor sufficient to fix a system so an accident doesn't happen again. It might act as as an incentive to fix a system, but it's less direct than actually working on fixing the system.

A starting note: I don't disagree with you (about systemic issues), but I want to explain what I understand as the perspective you are responding to.

A "scapegoat" is someone who is incorrectly blamed for someone else's errors or sins. The perspective you're responding to is this: They built the system, they run the system, they have continuously warned "This system is dangerous!", and yet persisted. That is not being incorrectly blamed, not being a scapegoat, and instead is a collaborator.

So I think you mean to ask: "Why do people focus so much on finding someone to blame?" It's not merely semantic, because the answer to that is more straightforward: Consistent accountability is a major factor in deterring bad behavior. It is not the only factor, but it is a major one.

That is my Steel Man understanding of the people searching for individual blame.


This is a good comment and I'd like to add that I sometimes am astonished how many comments on HN are on the topic of liability or blame for this or that. I see liability as a 2nd or even 3rd order phenomenon of most of the subjects under discussion, but for whatever reason, liability always seems to come up. Maybe because it's sort of easy to reason about? Thirst for justice? Natural reaction to feeling disempowered? Whatever it is, I'd like more focus on the actual topic.

Sometimes in "normalization of deviance" situations, there isn't anyone specifically to blame. I wouldn't make an assumption that you can find anyone particularly blameworthy without doing an investigation first.

The rate at which these "labs" are creating these incidents is simply staggering.

Imagine we made nuclear weapons a private industry, had CEOs bragging how they have enough warheads to blow the Earth to smithereens, and then they "accidentally" nuked three cities over a short period of time each, saying they lost control, or rather couldn't contain their semi-autonomous weapon. All somehow managing to turn the PR around from their abject incompetence and towards SciFi visions of mankind hunted by self-replicating bombs.


> Why do people focus so much on finding scapegoats?

I, for one, am not trying to find scapegoats or go on a witchhunt.

But managers are paid a lot of money to take responsibility. Yes, that's an old school thought, responsibility. But that's one big reason they get a big, fat paycheck.


Umm because it costs money to defend your companies servers when someone “accidentally” hacks them.

Countries demand reparation for damages in war. Citizens of those countries sue for damages and win.

Accountability is not a foreign concept. And the point is to disincentivize negligence. Because negligence is cheaper. And in this case, accidental hacks are marketing spend.


> Historically, many of these problems were bottlenecked by human attention. Someone had to care enough to spend hours or days reading obscure material, testing unpromising ideas, tracing references, and trying things that might go nowhere

I wonder how many of the recent results are due to the fact that very few looked at the problem to start with. Still great results, but the general impression is that it's more about the so many low-hanging fruits than the actual capability.


Let's not normalize the achievement. Just a couple years ago this would be considered science fiction. We can argue that 2026 AI can't solve the very toughest cryptograms, but the fact it can solve nontrivial ones is already magical.

Now on to the Voynich Manuscript :)


"AI solves niche thing you've never heard of" is a daily headline at this point. What's genuinely cool isn't that AI managed to solve some specific problem only a handful of people even cared about, it's that humanity can now cheaply clean up its backlog of such things.*

That does not mean that specific instances of it are still very interesting though. This article is the "I had claude vibecode a thermostat for my bathtub" of cryptography.

* And in this case I'm not sure it even meets that bar. For all we know a couple readers back when the book released had a delightful afternoon with it, solved the riddle, then forgot about it.



I disagree. This may be some niche thing that I've never heard of but a) it's still non-trivial; it still would have been science fiction to solve it a few years ago, and b) have you already forgotten the Navier Stokes drama? That is not some niche thing I've never heard of.

It kind of blows my mind how quickly people have forgotten both the state of AI in ~2010, and the outlook. If you had asked 100 people in 2010 whether they would see AI that could actually pass the Turing test in their lifetimes, you would have got 100 "no"s.

AI had been an unsolved problem for literally decades and it was firmly in the nuclear fusion/flying cars category.


> If you had asked 100 people in 2010 whether they would see AI that could actually pass the Turing test in their lifetimes, you would have got 100 "no"s.

There's no way that's accurate.

We already had big claims of the Turing test being passed in 2014, by a bot that had been doing almost as well for years.

There were plenty of people expecting fusion in their lifetimes too and that's going okay.


> We already had big claims of the Turing test being passed in 2014

Nah there was that bullshit Loebner prize or whatever, but that was just shitty chatbots being "judged" by people asking questions like "how are you today?" and then being breathlessly reported by the press. It was a publicity stunt.

Perhaps I should have said "100 people well-informed about AI".


From what I can find of early 2000s predictions, plenty of informed people saw good odds of strong AI by 2050.

"Yeah it is just token predictor, it could not even beat top 0.01% domain experts so it does not count as intelligence"

that's the point, it is clear the the bar to determine if llms are useful / intelligent is being moved every time these systems improve, but it is starting to fell like people are in denial. we are seeing significant progress, at a rate we are absolutely not used to experience.

They deny what: that they're very impressed, that they think it's cool, that their minds are blown, that they really really like it? Denying these things is allowed, you know.

Edit: if you're going to try to stage an intellectual wrestling match on the topic of is this thing awesome or not, you might as well make it a proper wrestling match and maybe get greased-up Turkish style. It would be more entertaining and you'd be more likely to arrive at a meaningful conclusion.


you forget the bar was moved very far down with all the slop we're experiencing with images, music and especially code

you can't claim "you keep moving the bar", if the very first thing when an LLM drooled out a piece of code, was to proclaim "this is good enough cause it gets the job done" followed by a barrage of "we're not quite there yet but exponentials or something, so very soon it'll be incredible"

yeah, from there it sure looks like "moving up the bar"


Im very impressed by the technology, I must admit that I didn't expect it to advance this fast.

But, at the same time, I'm really tired of there always being some shroud of dishonesty (ex. navier stokes and the two mathematicians working on it).

At this point my default is that I don't blindly trust the companies, I try to keep in mind they are trying to sell their product and win market share, there are so many perverse incentives at play I just can't take anything at face value.


Aren’t many of the greatest problems obscure to the lay?

I personally think that beyond normalizing, we should be actively be trying to dismiss this with all the cynicism we have. What does Anthropic have to gain from writing this? Behind the the scenes what might Anthropic be failing to disclose? How many failed experiments do we not know about?

The article doesn't seem to be written by Anthropic.

The Voynich theory I find most compelling is that it was a hoax made for a quack doctor, made to look like a foreign herbal manuscript. "Oh of course the local doctors can't help you, but my special book from a faraway land that only I can read may have the cure." Some recent analysis of the manuscript has found that the pages are more linguistically similar when read as individual flat sheets than how they're read when bound (i.e. whoever was writing the text was most likely going sheet by sheet and using the last completed page as reference for the text). The manuscript has only been bound once, in the fifteenth century (around when the vellum pages have been carbon-dated to), so whoever bound the manuscript was not able to "read" it. See https://journals.openedition.org/digitalmedievalist/2331

It is indeed absolutely incredible that it can solve these puzzles given plaintext instructions with very little context.

This was a basic cypher that effectively no one cared about. It wasn't famous or particularly notable.

I would say, the only reason it was never solved was because not enough people actually cared about it to begin with.

This isn't a big accomplishment.


Humans are incredibly good at adapting. A few days ago AI solved Navier-Stokes and I was blown away. Now I'm already thinking: "Well, it was only a counterexample and it brute-forced its way to it." lol

> AI solved Navier-Stokes and I was blown away.

That's not what happened, go read about it harder, please.


Sure. More precisely: they resolved the Navier-Stokes Millennium problem as posed by the Clay Institute. Not sure what else "solving Navier-Stokes" could reasonably mean. A general closed-form solution probably doesn't exist. And numerical solutions have existed for decades. But of course there are still open questions like unforced solutions etc.

I suspect EdwardDiego is referring to the brouhaha about whether OpenAI's training for the model that produced the alleged solution to the Millennium Problem about the Navier-Stokes equations was trained on material that included conversations Tristan Buckmaster and Levent Alpöge had had with earlier OpenAI systems.

I think there's a bit less to that than meets the eye. Yes, OpenAI's result builds on human work. It's possible that it builds on more human work than OpenAI admitted. But even if we suppose that everything Buckmaster and Alpöge did (which, btw, was itself very heavily LLM-assisted/generated work) was a necessary precursor to what OpenAI released, it's still the case that OpenAI's clankers completed the solution and Buckmaster and Alpöge didn't.

My understanding from what Buckmaster has written about this is that the deep mathematical ideas behind their work (and presumably OpenAI's) are due to Córdoba and Martínez-Zoroa. Those ideas are in the published literature, and human mathematicians and AI systems alike are allowed to use them, and doing so doesn't mean they didn't actually do something impressive. Mathematicians build on one another's work; that's how mathematics progresses and always has been.

It may very well be that OpenAI's announcement has a serious problem of professional ethics, especially as their first version of it didn't even list Córdoba and Martínez-Zoroa in its references. (On the specific question of what if anything they learned from B&A's work before that was published: OpenAI are now claiming that after investigating carefully they are confident that the model was not trained on anything Buckmaster and Alpöge did after early July. B&A had been working on this thing for much longer than that. However, on Buckmaster's account of things it wasn't until mid-August that they got beyond what he calls "preliminary results".)

But! The results of B&A were themselves largely AI-generated. (From Buckmaster's statement: "on August 15th, we obtained the blow up results, with smooth forcing, for both Boussinesq and Euler. I can say the first LLM generated proof Levent sent me was the most horrendous I have ever read; we verified it on Lean on August 22nd. Since this point, we have been working around the clock to understand this proof and turn it into something readable." That is: the LLMs found the proof, and B&A had to work to understand what the LLMs had done. It's not that humans did the thinking and AIs just did the gruntwork. (Except in so far as one might want to give all the credit for Real Deep Cleverness to C&MZ.)

And! What OpenAI say their model has proved goes well beyond what B&A did.

I don't see any way of slicing this that makes it unreasonable to say (unless it turns out that there's an error in the proof -- unlikely, given that it comes with Lean verification, but there have been misformalizations and Lean bugs in the past and there surely will be in the future) that AIs solved the N-S problem. No, they couldn't have done it without the work of C&MZ, but again: important mathematical work almost always builds on earlier important mathematical work, that's just how it is. Yes, if OpenAI are lying through their teeth their model might have had early access to B&A's ideas -- but it seems like most of the B&A work was actually done by AI systems anyway.

It is (I think -- I am not an expert and in particular I have not so much as looked at OpenAI's publication) reasonable to say that the deepest ideas here came from humans, and that it was already widely expected that the N-S problem would be solved in the not-impossibly-distant future in something like the way it has been. So, sure, what the AIs have done here is much less impressive than if they'd settled the Riemann Hypothesis or (probably even harder) PvNP. But it's still a resolution of a famous mathematical problem that any human mathematician would have been very proud to have achieved.


Someone wrote a prompt, that included instructions for finding the problem itself and got handed a solution by a machine trained on all available text. I don’t see any achievement for the prompter. As for the machine, we can’t keep being perpetually shocked 24x7. It’s tiring (unless if we’re being paid for it)

No, he's right. Actually, let's have a bit of sobriety when discussing the achievements of the most heavily marketed technology of all time, as published by an organisation that stands to benefit financially from the public perception of that technology. The discussion of "what made this problem low hanging fruit" is much more interesting, imo, than just breathlessly joining the hype train.

Thank you.

More money than the GDP 90% of the sovereign countries around the world is hanging in the balance, and people are taking everything OpenAI and Anthropic are saying at face value as if this isn't the financial / marketing equivalent of war, assuming they they wouldn't use every legal and shady tactic, bending every truth available to them to sway the balance of public opinion in their favor. It makes me feel like I'm living in the twilight zone. People need to wake up.


A few months ago a post claiming an amateur deciphered Linear A hit the front page and quickly got almost 500 upvotes:

https://news.ycombinator.com/item?id=48600107

https://aiclambake.com/clamtakes/linear-a/

Despite the announcement originating from a blog named "AI Clambake" covering "weekly, human-powered newsletter for advertising folks". Written by a personal friend of the author. Announced without any corroboration or commentary whatsoever from academics or subject matter experts of any kind. And, of course, not submitted to any peer reviewed journal or even Arxiv.

The author of the purported discovery was described as a "self taught AI engineer and amateur linguist". In the comments the friend insisted several times that a draft of the paper (not posted), was emailed to a top professor at Rutgers, giving it additional credibility that his friend wasn't another one of ten thousand cranks who has made the same claim over the years (seemingly unaware that cold emailing random professors found from a Google search is the first thing basically every crank does).

You would think this should have set of dozens of alarm bells for everyone, making the value of this announcement basically zero. And yet it hit the front page with the bulk of comments ecstatic that some random guy with Claude Code could do something experts in academia who spent their lives devoted to the problem couldn't.

I had become accustomed to the toxic optimism of this hype cycle in which even mild criticism leads to accusations of being a discredited "AI skeptic"/Gary Marcus/Ed Zitron type who was "coping" (?). But this was like something you'd see shared on FB linking to a .xyz domain by an elderly family member who recently drained their accounts buying Xbox gift cards to pay their IRS bill.

It feels a lot like the week or two when HN was overflowing with exuberance from the LK-99 room temperature superconductor "discovery ". You'd see post after post fantasizing about an imminent future with a world full of maglev hovercrafts, MRIs built into every phone, fusion reactors and more. But people pointing out none of that was scientifically plausible and evidence of LK-99 superconductoring was non-existent were accused of knee-jerk negativity and the typical HN cynicism and pessimism.


There's nothing AI-specific about this. It's routine for all decipherment claims.

Compare e.g. https://arstechnica.com/science/2019/05/no-someone-hasnt-cra... . (It's a debunking, but the reason a debunking got published is the media hype frenzy beforehand.)


Wait we don't like xyz domains??? :(((( !!?

I’m pretty sure the Beale ciphers are a hoax, but I’d love to be proven wrong.

What's so magical about the problem... Its the exact time of problem they were built to solve (things that can be brute forced with language). I'm not impressed.

Why don’t you solve such problems?Being “not impressed” sounds more like a knowledge gap on your part than an informed opinion.

Pepper grinders have always impressed me, very effective devices, great at grinding out results, I mean pepper. I can't do it by hand at all!

Sure you can, its just difficult to the point of being infeasible

Why are you so toxic?

How ironic for you to say that to me. Maybe you can use ai to help explain it to you.

It's not particularly surprising that if you:

(1) take a 300 line NN algorithm,

(2) throw a quarter of the world's GDP + all literature ever collected at using the algo to train a NN

(3) throw another quarter of the world's GDP at billions of teraflops for inference, and

(4) aim the resulting world's-largest-computer at marketing itself to investors, for instance by decoding ciphers from obscure medieval manuscripts,

that you could perform some pretty magical tricks. There are other feats humans have performed for less cost, like sending people to the moon, or landing a rocket vertically, or idk, curing Polio.

I'm not knocking the "miraculous" advance here. The unique thing about the solution which makes it particularly non-trivial and something that humans would struggle with was exactly what LLMs excel at: Diffing loads of texts against each other. But the 176k tokens at around $10 doesn't tell the story of the cost. It says a lot about the externalized cost and the amount of money flowing in to support the hardware. If they'd put a $100,000 bounty out to solve that cipher, I think the internet would've solved it in a couple days.


It's funny we're already at the "actually this isn't very impressive" stage when it was a little over a year ago when we were making fun of LLMs for not being able to add numbers.

IC production takes a vast amount of resources and wealth, and it's a known quantity (after all, we've been doing it for decades), but it's still impressive what modern fabs can achieve.


Afaik LLMs still can't add numbers. They've just had their system prompts updated to make them use a tool call for any arithmetic.

That does not reflect with testing I just did locally. When the option of writing and running scripts was available, Qwen3.8-flash did indeed prefer to just do a physics related calculation via code. But, with a fresh chat and tool access turned off, it did the work on its own and produced correct results. Each step was rounded similar to how a person working by hand might do things, but matched my own hand calculations perfectly.

I found this post hilarious exactly about this the other day:

First, it’s AI can’t multiply 4-digit numbers.

Then it’s AI can only, by brute force, get silver in the IMO with specialized systems.

Then it’s OK, well, now a general-purpose model can get gold, but it’s still just the IMO, it’s for high schoolers.

Then it’s OK, it can solve a few trivial Erdős problems, but only because nobody seriously tried them before, they were low-hanging fruit.

Then it’s OK, a lot of serious mathematicians tried this one, but the result was still obvious in hindsight, it just combined knowledge from a thought-to-be-unrelated field, if any human knew that, they would solve it.

And then to OK, but there are still Millennium Prize Problems.

Then OK well it's just Navier-Stokes wake me up when its the Riemann Hypothesis.

Then-


If AI does solve the Riemann it will be the watershed moment when the world realises what has happened.

The ability of the algorithm to absorb billions of dollars of training effort is itself the major breakthrough of the transformer architecture.

also fire is not particularly surprising, once discovered

Yes, even many of the proofs seem to be extremely long and complicated. The Navier-Stokes proof is 57 pages of very dense math and a pretty crazy amount of code: https://github.com/openai/NavierStokesAndEuler/tree/main/Nav...

Given the close relationship between compression and intelligence, I'm somewhat surprised at how poorly the cutting edge models do with being concise.


You know, the first time you navigate somewhere (if you don't already have perfect directions) will probably be the longest route you'll ever take to get there

For Earth, the proof presented for NS is just our first attempt navigating from our previously known facts to the proof.

I expect we will be able to shorten it dramatically (most likely with human and AI insights), but I don't think we should read too much into the length. If you want a similar point of comparison, see the original proof (by humans) of Fermat's last theorem. It has been shortened significantly. This is normal.


> The Navier-Stokes proof is 57 pages of very dense math

> how poorly the cutting edge models do with being concise

LLMs solve a Millennium prize problem. People complain the proof is too long, within a week. What a time to be alive!


Deep commentary on an unverified proof of this level requires extreme expertise. So the existence of some shallow commentary is uninteresting; it doesn't actually imply pettiness or deflection or anything like that because it's so hard to make your commentary any deeper at this time. And no we shouldn't expect people to say nothing.

My observation isn't that solving a Millennium prize problem is not impressive.

It is the mechanism the LLMs use to do it. They seem to accel right now at quantity of work over quality of work. I'd be willing to wager there is a much simpler way to achieve the proof.

I've also seen this with code, LLMs do get the job done, but they tend to write 10-100x more code than humans to get the same job done. Still a massive value gain because they can write that much code extremely quickly.


"if llms are so intelligent why is cancer still not cured?"

I'm waiting for the inevitable: "Well, LLMs only cured a few types of cancer"

Then: "LLMs only cured cancer as a marketing stunt."

>I'm somewhat surprised at how poorly the cutting edge models do with being concise.

because they're not intelligent in the sense you're hinting at (conceptual integrity or generalization) but they are as the name suggests, large. Like comparing a forklift to a human. It's easier to bulldoze through a lot of things than tie your shoes.

If we weren't quite as impoverished conceptually and still had the vocabulary of the Catholics we'd recognize this as ratio (discursive knowledge) vs Intellectus (apprehending knowledge)


What an incredibly useless comment. You state a conclusion as fact without any supportive reasoning/evidence.

Prove that human intellect is different and that we solve problems using fundamentally different processes. I’m waiting.


>You state a conclusion as fact without any supportive reasoning/evidence.

No, it's the other way around, it's a reductive view on intelligence that mistakes its own methodology for ontology.

It's obvious to see that there's no intellect in an LLM as defined above because of how they work. LLMs put one token in front of the other, they don't work towards formal ends, there's no intentionality in them. They don't synthesize the information they process into a unified experience. Thinking an LLM can apprehend what it does because it can process large amounts of text is like thinking your TI-83 understands math because it can multiply large numbers.

That's also why the failure modes of LLMs are what they are. They can churn out tens of thousands of lines of code but also just as easily go in circles like a roomba. They can process an entire encyclopedia but not solve problems a 10 year old can solve.


What's an example of a problem that an LLM can't solve but a 10 year old can?

"Prove"? Like, a formal proof, about intelligence?

yeah you know, that concept we've never been able to define using language, making heavy use of the human experience which can also not be captured in language (proof: how bad LLMs are at poetry)

the question is, when comparing a human and a large language model, whether the intellect (that cannot be captured in language) is different from anything the language model can actually do (e.g. language)

the answer to this seems quite obvious to me, and I would actually posit that the onus is on the other side, to prove they are even remotely similar

maybe people think that the voice in their heads is what is doing the thinking? is that the confusion here?


Do you think the LLM is the Chain of Thought? Did you also get confused by the name? Because, much like humans, the CoT is a tool to narrativize and maintain internal coherence. The actual thinking happens invisibly, in the forward pass. Just like...

I wonder though how much of science has similar issues, that there are hundreds of semi-promising but niche areas that require tons ton of deep analysis that would have simply cost too much to explore all areas, but now become feasible.

The Navier-Stokes proof cost ~10mio USD in tokens at consumer prices, and it was something that the two people working on it were allegedly weeks or months away from solving. You can buy ~200 Math PhD Years for 10Mio USD, so it doesn't seem like a shortcut at all (actually it sounds like we would have gotten the same solution for ~1-0.1% of the price if we had just been a little patient.

We weren't willing to pay for 200 math PhD students to try to find singularities in Navier-Stokes, I am skeptical of how much we would be willing to pay OpenAI to do research on "niche scientific areas"?


Which is why some automatic intelligence is so valuable

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: