Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Something similar is also described here: https://docs.anthropic.com/claude/docs/claude-says-it-cant-d...

> This can be a way of getting Claude to comply with tasks it otherwise won’t complete, e.g. if the model will by default say “I don’t know how to do that” then a mini dialogue at the beginning where the model agrees to do the thing can help get around this.

This “vulnerability” definitely isn’t new, I’d even say it’s obvious to anyone who understands how LLMs work



The paper makes it clear that it's building on past work, and that the novel part of their method is to automate the process, and the interesting result here was that the suffixes were transferrable.


To be honest I didn’t actually read it and just looked at the title (which seems to have been changed now)




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: