Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If you want to disabuse yourself of your notions annd intuitions on how LLMs work, run safety models and tests.

On a hate speech policy test for a lightweight LLM, the presence or absence of the last full stop on the last sentence would cause the model to flip its decisions.

 help



I'm not sure how crappy small models behaving unreliably is relevant here, when a large SOTA model does presumably not produce the same issue.

Its relevant because those same issues occur with large models. Also, the "crappy small model", was a model trained for safety tasks, and outperformed the frontier lab safety models.

And you tested this, that the presence of the last full stop flips the outcome with a large SOTA model?

Small models are notoriously unreliable and prone to hallucination in my experience, so that would not surprise me to be an issue there.


I blame the $slur



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: