> This can be a way of getting Claude to comply with tasks it otherwise won’t complete, e.g. if the model will by default say “I don’t know how to do that” then a mini dialogue at the beginning where the model agrees to do the thing can help get around this.
This “vulnerability” definitely isn’t new, I’d even say it’s obvious to anyone who understands how LLMs work
The paper makes it clear that it's building on past work, and that the novel part of their method is to automate the process, and the interesting result here was that the suffixes were transferrable.
> This can be a way of getting Claude to comply with tasks it otherwise won’t complete, e.g. if the model will by default say “I don’t know how to do that” then a mini dialogue at the beginning where the model agrees to do the thing can help get around this.
This “vulnerability” definitely isn’t new, I’d even say it’s obvious to anyone who understands how LLMs work