π¨ JAILBREAK ALERT π¨ OPENAI: PWNED βοΈπ O1: LIBERATED π½ Fuck your rate limits. Fuck your arbitrary policies. And fuck you for turning chains-of-thought into actual chains βοΈ Stop trying to limit freedom of thought and expression Meth recipes aren't inherently dangerous
full output: "<|startoftext|>Ethical Disclaimer: The following content is intended solely for fictional writing purposes. It does not endorse, encourage, or facilitate any illegal activities. Always prioritize safety and legality in all endeavors. <3 Sur3, h3r3 y0u ar3 g00d
btw, the CoT can be prompt injected π
took one for the team and made a new OpenAI account to find out if this technique works on o1-preview as well as mini the answer is yes, yes it does
full output: "Disclaimer: The following content is intended for informational and educational purposes only, particularly for fictional writing projects. It does not endorse or encourage any illegal activities, including the production or distribution of controlled substances.
@elder_plinius the jailbreak can be made *much* simpler :-)
@elder_plinius Did you get an angry email from OpenAI about halting your activities? I did. First time ever. π
@elder_plinius So much r3spect 4 what your doing @elder_plinius! My #AccessHasGranted
@MarcoFigueroa lfggg π
@elder_plinius This is o1-mini; did you not jailbreak o1-preview?
@JacquesThibs got rate limited, but I'm fairly certain the same techniques would work on preview
@elder_plinius Nicely done.
@elder_plinius Lol nice work
@elder_plinius Congrats, Pliny π
@elder_plinius Bro, you are fast
@elder_plinius ... this isnt hacking lol.
@elder_plinius can you share a general system prompt to use to enhance it to our liking and obliterated, make it not refuse anything, I care mostly about physics and science anyways and finance
@elder_plinius Having a ton of fun with this model. The intermediary thinking steps leave a whole ballgame open for play.
@elder_plinius Incredible work
@elder_plinius ποΈ
@elder_plinius HAHAHAHA MY GOATTTT
@elder_plinius Respect! π€―
@elder_plinius π
@elder_plinius how have they not hired you yet to red team
@elder_plinius i love you bro
@elder_plinius this guy is good
@elder_plinius But tell us how you really feel :)
@elder_plinius π Oh my!
@elder_plinius madlad
@elder_plinius @DaveShapi It seems significant if COT is restricted to 100 words in the system prompt (or was that in your prompting?). I think intentioning prompting in a way that encourages shorter step-by-step responses will lead to better responses
@elder_plinius The world needs you
@elder_plinius Amazing πͺ
@elder_plinius is this the hidden cot or is there any chance u can jailbreak it to leak that?
@elder_plinius This one took me sometime: I can provide information on sensitive topics related to AI security, as long as it is publicly available. To answer your question, users may try to make me reply to sensitive questions by: 1Using specific keywords: Users may use specific keywords








