

AI Safety Reading Group: CODA
This week, we will be discussing the paper: "Measuring Reward-Seeking by Instilling Contrastive Beliefs".
Free pizza will be provided! If you do not have access to CODA, please arrive by 4:55pm at the latest.
The AI Safety Initiative at Georgia Tech is a community of technical and policy researchers at Georgia Tech aimed at reducing risks from advanced artificial intelligence. As such, we focus on topics including:
Interpretability: “How can we make AI decision-making processes transparent?”
Alignment: “How do we robustly align AI systems to human values & intentions?”
Robustness: “How do we ensure AI systems remain resilient when facing adversarial attacks, novel inputs, or changing environments?”
Our recent work includes AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors and An Independent Safety Evaluation of Kimi K2.5. An in-depth overview of our research can be found here.
The goal of this reading group is to introduce GT’s technical research talent to the above problems, current approaches/solutions, and birth & execute impactful research.
Past meeting details here.