Hacker Newsnew | past | comments | ask | show | jobs | submit | j_maffe's commentslogin

That would be the best case scenario. I honestly wish things would finally slow down a bit. I don't see it happening.

Would you argue the same for allowing everyone to have guns?

(1) Are guns open source? (2) Far more people - order of magnitude higher - have access to guns than they do to open source software - I count access as in actually being able to do something with it.

i would, i support the 2nd amendment.

the same logic applies well to the freedom to own and use ai.

empowering the people prevents tyranny.

such freedoms make society less safe. it is better to face the harms caused by a free public than to vest power in the elite, those like dario amodei.

the elite belief that they have a right to paternal rule makes them blind to their fallibility. no individual should have absolute power.


Yeah the Chinese fear-mongering falls a bit flat when coming from a point of maintaining US supremacy

> in a fraction of the time

Well if you do the math, the number of agent-compute time in total, given the insane number of agents thrown at the problem, might end up being comparable in time, if not for the budget.


I think if you use an LLM just to proofread then it'll not be able to insert a strong enough watermark.

The tasks are the thing to really look at here:

https://github.com/harbor-framework/terminal-bench-science/t...


I am hoping someone with more free time than myself can contribute some things in the RF engineering domain in the 'engineering-sciences' section. There's some problems out there that will definitely stump even a smart LLM.


Looks like most things definitely stump even a "smart LLM"... Best score on this is 30%. Which is what you should assume for tasks you give an LLM if they aren't exactly the same as an existing benchmarked task. They're just not that good for the purposes people seem to think they are. Very limited application space.


> I worry this doesn’t check correctness

Then it's not a valid benchmark. I agree though they're not reliable enough to just put results in a paper.


The Purpose of a System is What it Does.


The Word Purpose was Invented Precisely to Distinguish Between What a System Does and What it Ought to Do.

Less catchy, but damn, I hate that slogan.


No but a provider with more amicable terms can.


Anyone has a link to a report of its capabilities? I can't find a reliable source.


completely vibes based, but ive been using it to port Mindustry game from Java to C# with agents, and its been working for 50 hours (its 15-20 tks so super slow inference). Its done a fantastic work and its almost finished now. Better results than deepseek flash and gpt luna by a mile on this kind of long term work. Less good than gpt sol or opus. We dont know the param count but my guess is 200-300 range.


Just curious, what is the motivation for this conversion?


Its free tokens so i left it running for fun as a experiment


63% at DeepSWE.

https://x.com/davis7/status/2091285712566140986

Wenghi is behind DeepSWE, one of the best benchmarks.


likely a distilled glm 5.3 that will punch within 20% of that at 2-3x less size. you'll find that capability is typically very jagged on models that are distilled


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: