

Lets reach it together, by building data-driven AI engineering features live during our sessions.
#7 - Agent Evaluation
Bring your research mindset into AI!
This time we're going to tackle agent trace evaluation. Lets say you receive a long user-agent conversation, how exactly do you go about evaluating it? What is a good conversation?
As always, it's a live coding session. I'm going to pull a dataset and start hacking at it live, together with you.
Our sessions are friendly to all experience levels, on purpose, you are welcome to ask along during the zoom / chat. We won't get into deep theory but rather focus on practical ways to deal with agent evaluation.
Supporting methods: LLM as a judge
------- Links -------
Join build guild Whatsapp community (Hebrew)
Follow Substack posts for the long form pieces
https://substack.com/@mlarchitect
Or if you rather watch coding session in YouTube
Lets reach it together, by building data-driven AI engineering features live during our sessions.