AI Safety workshop
What: Audra Zook will lead us through an LLM jailbreaking workshop
Frontier models aren’t supposed to tell you how to build a bomb, make meth, and poison your boss, but they will with the right input. Welcome to the world of LLM jailbreaking, where all sorts of dangerous advice* is only a prompt away! Join competitive jailbreaker Audra to explore this AI safety risk posed by existing LLMs.
This event will have two parts. First, Audra will give a brief talk on common jailbreaking techniques, current safeguards to prevent jailbreaks, and the evolving LLM red-teaming landscape. Then, she will lead a workshop in which participants will develop and test their own jailbreaks (in an approved environment–no one will be using their personal LLM accounts). Bring your laptop if you can! Where: The Hub, The Primary (3rd Floor), 26 Broadway 3rd floor, New York, NY 10004 When: Thursday, January 29th 7-9pm