

The One About Evals
How do you know if your AI is really working? It's one thing to build AI systems - it's another to be confident they're doing what they should. Join us for practical insights from three different perspectives on AI evaluation and safety testing, as we explore how organizations across education, recruitment, and government are tackling this crucial challenge. 🧐
Watson Chua (Lead Data Scientist) and Shaun Khoo (Data Scientist) from GovTech's AI Practice team delve into the comprehensive evaluation and safety testing of AI models in educational chatbots. Watson will share insights on developing systematic benchmarking approaches for assessing AI model performance, whilst Shaun will explore the framework established for safety testing and continuous improvement. Together, they'll present valuable lessons on optimising AI capabilities whilst maintaining robust safety standards throughout an AI product's lifecycle. This session offers practical insights for those interested in the responsible development and deployment of AI in educational settings.
Sreejith Balakrishnan (Senior Responsible AI Scientist) from Resaro, will share an evaluation use-case from the recruitment industry focused on Large Language Models (LLMs). The session will highlight key business risks identified, the evaluation methodologies employed, and lessons learned. This sharing offers a glimpse into the types of AI assessments conducted by Resaro, while shedding light on the practical challenges of evaluating LLM-based applications in real-world settings.
Tan Wen Rui (Senior Manager, AI Governance and Safety) from IMDA will be sharing more about the AI Verify Foundation’s Global AI Assurance Pilot. The sharing will dive briefly into the use cases and the next steps of the initiative.
About the speakers
Watson and Shaun lead teams in Search & Retrieval and Responsible AI respectively, in the AI Practice team under the Government Technology Office (GTO) of GovTech. Their teams focus on exploring new technologies and distilling best practices for AI systems in their respective tracks.
Sreejith Balakrishnan is a Senior Responsible AI Scientist at Resaro, where he leads the evaluation of AI systems across various domains such as healthcare, recruitment, and facial recognition. He specializes in trustworthy AI, including performance, robustness, and fairness assessments of Large Language Models and Computer Vision Models. Sreejith holds a Ph.D. in Human-centric AI from the National University of Singapore and has represented Singapore in global AI standardization efforts.
Wen Rui is a Senior Manager with IMDA's AI Governance and Safety team, where she develops policy guidance and tools for responsible AI implementation. She co-authored IMDA's Model AI Governance Framework and related guides, helping organisations adopt AI responsibly. In 2023, she was instrumental in establishing the AI Verify Foundation, an open-source initiative for AI testing and trust.
NOTE
Session will begin at 3pm
Lorong AI members will be prioritised for all Lorong AI programmes
Want to be part of our community sharing sessions? Sign up as a speaker here!
———
About Lorong AI
Lorong AI is a co-working hub where AI practitioners connect, share knowledge, and grow through curated programming and a collaborative environment. Home to programmes like AI Wednesdays, AI ToolsDays, Fri-DIYs and more, Lorong AI offers hands-on workshops, technical deep dives, and opportunities to collaborate and solve real-world challenge with AI.