

AI Evals Your Whole Team Can Run: Lessons from Nova Escola
Most organizations building AI products face the same bottleneck: evaluation lives with one or two experts, while the people who best understand the users (educators, health workers, program staff) stay on the sidelines. Yet these domain experts are often exactly who you need to judge whether an AI's outputs are actually good.
In this webinar, Lucas Rocha, AI Product Manager at Nova Escola, will share how one of Brazil's largest teacher platforms evaluates its WhatsApp-based AI lesson-planning assistant, which supports hundreds of thousands of public school teachers. Walking through the four-level AI evaluation framework, Lucas will unpack the questions his team asks and answers at each level: At Level 1, does the AI produce pedagogically sound lesson plans? At Level 2, are teachers adopting the tool and coming back? At Level 3, is the tool changing how teachers plan and teach? And at Level 4: does all of this ultimately improve student learning?
Specifically at Level 1 (AI model and system evaluation), Lucas will also share how Nova Escola turned it from an expert-only activity into a practice the whole team runs. Lucas will show the workflows they built to bring pedagogy specialists and product managers into the evaluation loop: expert-designed rubrics, LLM-as-judge pipelines aligned with human judgment, and in-house review tools and reusable AI workflows that let non-technical colleagues run error analysis and labeling alongside engineers.
The webinar will include a presentation followed by a moderated discussion and audience Q&A.
About the speaker:
Lucas Machado Rocha is an AI product manager at Nova Escola, a Brazilian non-profit that supports public school teachers. He has spent over 15 years connecting product, technology, and social impact. At Nova Escola, he builds and evaluates an LLM-powered teaching assistant that reaches teachers over WhatsApp, and he has been teaching evals to other product managers across Brazil.