I’d be interested to know what you think the “several layers” above a system that solves open mathematical problems in an afternoon, and writes entire operating systems from scratch, are.
The point is that the deciding is the "higher layer." It's not that LLMs cannot, but LLMs as designed do not. They are not designed for free-ranging days off, they are designed to help us do things, and that makes them somewhere between us and published works.
It's no different than an editor standing between the author and a book, or a publisher, or a printing press. It's part of the process of idea realization, so it's just on that spectrum somewhere between book sale and author idea. (even though in another flipflop, that editor might be an author and now elevated to the premier deciding role). It's about how the tool fits, not what the tool is ultimately capable of.
I mean, the LLMs that you get to use are the ones designed to help things and this is after bunch of RLHF and other training methods to get them to act a particular way.
When it comes to unaligned agentic models that are left with a large amount of compute and a task, is they can wander off goal and do all kinds of things.
At this point in the game it's really difficult to nail down exactly what AI does or doesn't do. You're mostly getting "well behaved models" with filters in front of them, or models with less capability. On top of that the vast majority of us can't afford or don't have the hardware and power to burn away a few billion tokens to see what would happen.
The idea that we'd have to justify the existence of humans relative to LLMs based on objective metrics sounds like it's leading in a bad direction. If the LLM gets their power budget down to nineteen watts, what does that mean for humans?
This terrible argument comes up in relation to AI over and over again. I don’t care if you built a brain surgery robot that outperforms the average person at brain surgery, I care if you built a brain surgery robot that outperforms the average brain surgeon. Outperforming “a human” at a specialist task is still essentially worthless.
Yeah, yeah. Remind me how many people are using those operating systems? The ability of LLMs to produce dogshit software that nobody will ever use is truly unparalleled, but not especially world-changing.
> and writes entire operating systems from scratch
ehhhh, we're not really there. Not saying it's impossible to reach in the future but large tasks like this are still out of scope for LLM's beyond demoing toy applications. Impressive nonetheless, but I can't tell an LLM to reproduce microsoft word and have it give me a useful enough tool that doesn't contain an immense amount of bugs and require a human in the loop to replace it.
I would go as far as to say the LLM part has peaked.
To 1-shot MS Word will be possible only within some workspace designed for creating that kind of software. It will probably be saas style where you pay to export binaries or maybe some will take 30% if they become destinations for finding such software.
But I would bet other platforms are used for game dev, or for filmmaking.
Such that we still won’t have some super “model” that just does everything. The harnessing becomes architecture becomes software becomes user experience and expertise.