Fine-Tuning the Agent Stack: Harnesses, Evals, and What It Takes to Reach Production

Register to See Address
London, United Kingdom
Registration
Past Event
Please click on the button below to join the waitlist. You will be notified if additional spots become available.
About Event

As AI agents move beyond demos and into real-world workflows, performance depends on far more than the underlying model. The surrounding harness, tool access, feedback loops, and evaluation systems increasingly determine whether an agent is reliable enough to deploy.

Join Prolific and AI Circle in London for an intimate conversation with researchers, engineers, and technical leaders exploring how teams are fine-tuning agentic systems and evaluating them under real-world conditions.

The discussion will examine how to design stronger agent harnesses, build evaluations that capture long-horizon behavior, identify meaningful failure modes, and use high-quality human feedback to close the gap between promising prototypes and dependable production systems.

Expect a technical, candid conversation among peers, Prolific, and AI Circle community members, followed by drinks light bites, and networking with the people actively building and evaluating the next generation of AI agents.

Moderator:
Joséphine Parquet - London Chair at AI Circle, where she helps grow a community of AI researchers, founders, and operators. She's currently leading API product at Veed, building video avatars, and hosts Field Notes, a podcast of conversations with AI researchers.

Panelists:
Tyler Edwards - CEO and Co-Founder of Overmind, which automates the whole path from production agent traces to a shipped open-weights model. He spent 6 years building ML systems for British Intelligence before founding the company.

Thomas Mann - Thomas is a Research Engineer at Meta working across AI, software engineering, data, and cloud infrastructure. He is also the creator of ClawBench, a benchmark exploring how effectively AI agents can complete everyday online tasks, with a particular interest in real-world agent evaluation and reliability.

Angelos Perivolaropoulos - Angelos leads speech-to-text research engineering at ElevenLabs, working across research, inference, backend systems, and infrastructure. His work includes ElevenLabs’ Scribe speech-to-text models and research focused on making voice AI accurate and reliable in real-world conversational environments. 

Agenda

  • 6:00 PM - 6:45 PM Check in, Drinks & Networking

  • 7:00 PM- 8:00 PM Expert Panel & Spotlight

  • 8:00- 9:00 PM Food & Informal Breakouts
    ____________________________________

    About the hosts:
    Prolific powers high-quality human data for model evaluation and training. Learn more about Prolific.

    AI Circle curates high-agency operators across research, startups, and enterprise. Learn more about AI Circle.

Location
Please register to see the exact location of this event.
London, United Kingdom