Cover Image for #7 - Agent Evaluation
Cover Image for #7 - Agent Evaluation
Avatar for BuildGuild
Presented by
BuildGuild
The road to mastery crosses 1000 hours.
Lets reach it together, by building data-driven AI engineering features live during our sessions.
Hosted By
44 Going

#7 - Agent Evaluation

Virtual
Registration
Welcome! To join the event, please register below.
About Event

Bring your research mindset into AI!

This time we're going to tackle agent trace evaluation. Lets say you receive a long user-agent conversation, how exactly do you go about evaluating it? What is a good conversation?

As always, it's a live coding session. I'm going to pull a dataset and start hacking at it live, together with you.

Our sessions are friendly to all experience levels, on purpose, you are welcome to ask along during the zoom / chat. We won't get into deep theory but rather focus on practical ways to deal with agent evaluation.

Supporting methods: LLM as a judge

------- Links -------

Join build guild Whatsapp community (Hebrew)

Open Google Form

Follow Substack posts for the long form pieces

https://substack.com/@mlarchitect

Or if you rather watch coding session in YouTube

serjsmor

Avatar for BuildGuild
Presented by
BuildGuild
The road to mastery crosses 1000 hours.
Lets reach it together, by building data-driven AI engineering features live during our sessions.
Hosted By
44 Going