

AI Safety Poland Reading Club #6
Hosted by Piotr Kędziora & Kacper Dudzic
Registration
Past Event
About Event
Who should come: Anyone interested in AI safety, machine learning research, or the broader societal impacts of AI systems.
What it's about: We'll read and discuss influential papers in AI safety research. Read the paper beforehand, show up, and talk through the ideas, questions, and disagreements it raises.
Current agenda: We will be continuing the Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback paper, this time going through the 3. Preference Modeling for Helpfulness and Harmlessness and 4. Reinforcement Learning from Human Feedback sections.