

Subscribe to this Calendar for Event Updates
VAM! AI Reading Group: How to Correctly Report LLM-as-a-Judge Evaluations
📄 Paper: How to Correctly Report LLM-as-a-Judge Evaluations
Zoom link: https://servicenow.zoom.us/j/99425770821
https://arxiv.org/html/2511.21140v1
Presented by Helena Vallée
To maximize engagement, please try to read the paper in advance.
Summary: Many benchmarks nowadays are using LLM-based evaluation because it is the best thing we have. LLMs though tend to introduce noise in their evaluation. This paper discusses some good practices on how to calibrate such LLMs.
Want to present?
The list of papers will be available here: https://docs.google.com/spreadsheets/d/1HET5sjnHjwiF3IaCTipR_ZWspfgglqwdBFRWAfKBhp8/edit?usp=sharing
To connect with the group, join the Discord: https://discord.gg/teJvEejs94
Please arrive by 6:35pm the latest, because I won’t have my phone available to let people in afterward.
Timeline:
🕠 6:30 PM – Arrival & Networking.
🗣️ 6:45 PM ~ 7:15 – Paper Presentation
🗣️ 7:15 PM - Discussions
About the Facilitator
Issam Laradji is a Research Scientist at ServiceNow and an Adjunct Professor at University of British Columbia. He holds a PhD in Computer Science and a PhD from the University of British Columbia, and his research interests include natural language processing, computer vision, and large-scale optimization.
Looking forward to discussing the latest AI Papers!
Subscribe to this Calendar for Event Updates