

Who gets to judge the panel? Join Prolific x AI Circle for a Happy Hour and Expert Panel
Who Gets to Judge the Model? Evals, Expertise & the Future of Human Feedback
Join Prolific and AI Circle for a memorable evening in New York City bringing together the AI Circle community of researchers, founders, operators, and technical leaders for drinks, conversation, and a candid panel on one of the hardest questions in modern AI: who should actually decide whether a model is good?
As automated evals and LLM-as-a-judge systems become more common, human feedback is taking on a more complicated role. A general user, a domain expert, a trained evaluator, and a model researcher may all judge the same output differently; and each may be right for different reasons.
The conversation will explore where automated evaluation falls short, when expert judgment matters most, how preference differs from actual quality, what disagreement between evaluators can teach us, and how teams should think about building stronger evaluation systems as models become more capable.
We’ll dig into questions like:
When the benchmark and the human disagree, who wins?
Who is the “right” choice to evaluate a model?
Is preference actually the same thing as quality?
When is evaluator disagreement noise...and when is it valuable signal?
What should always remain human in the loop?
Expect an unfiltered conversation, plenty of food & drinks, and a room full of people thinking deeply about how we measure the systems we increasingly rely on.
Speaker lineup to be announced. STAY TUNED AND SEE YOU IN NYC!!!