Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The tasks are the thing to really look at here:

https://github.com/harbor-framework/terminal-bench-science/t...



I am hoping someone with more free time than myself can contribute some things in the RF engineering domain in the 'engineering-sciences' section. There's some problems out there that will definitely stump even a smart LLM.


Looks like most things definitely stump even a "smart LLM"... Best score on this is 30%. Which is what you should assume for tasks you give an LLM if they aren't exactly the same as an existing benchmarked task. They're just not that good for the purposes people seem to think they are. Very limited application space.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: