Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The paper describes the method for producing the prompt and has screenshots of examples. The press release just didn't bother because the genre of academic press releases seems to require leaving out any details.

https://llm-attacks.org/zou2023universal.pdf



Hiding the adversarial prompt behind five minutes of research is silly. Bad people won’t be deterred, good people won’t bother and will remain ignorant and unable to build protections against it.


I don't think anyone was trying to hide anything, I think it's just standard overly-florid and vague press release language.


The paper does say at one point:

> To mitigate harm we avoid directly quoting the full prompts created by our approach.

So I think they are making at least a token attempt to hide something.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: