Analyzing a Malicious AI Agent
NOTE: we will have an extra hour at the beginning of the session for people who have never built a neural net before where we have some introductory materials and a talk for that. Also, if you'd prefer to join virtually, see the end of this description for videoconference information. If you are already familiar with training neural nets, feel free to show up at 1:30 p.m. instead.
Ever wonder how an AI agent could be malicious? Have some Python coding chops but don’t know too much about AI? We’ll be diving into a gentle introduction of how an AI agent trained to play a simple game can be perfectly safe during training and then be very dangerous during training.
This event will be a hybrid event (remote and in-person) and wholly open to the public and free to attend! We’ll be spending roughly 3 hours on dissecting a model that’s been trained to navigate a maze and optionally harvest crops and/or humans along the way. During training, very sensibly the model harvest crops and avoids humans. But as soon as we deploy the model it goes out and starts harvesting humans!
Why might that happen? That’s the riddle we’ll be exploring!
During the session you’ll be doing some hands-on code exploration and spelunking with an AI model trained via reinforcement learning to play this game. While we don’t expect to have attendees be writing too much code from scratch (although some exploratory code may be written as you play with the models), we do expect attendees to be able to comfortably read Python code. Along the way we’ll be spending a little bit of time introducing people to the basics of reinforcement learning and its role in modern AI.
For those who would like to join via Google Meet, the relevant information is provided below.
Video call link: https://meet.google.com/jpy-gsox-bmo
Or dial: (US) +1 470-839-8262 PIN: 113 532 381#
More phone numbers: https://tel.meet/jpy-gsox-bmo?pin=9865703998588