

Post-Training, Evals & Self-Improving Agents - AAIF London x Prolific
Post-Training, Evals & Self-Improving Agents - AAIF London x Prolific
How do we make agents meaningfully better after the base model has been trained?
Join AAIF Community London, in partnership with Prolific, for a technical evening on post-training, agent evals, reinforcement learning, context, agent architecture and self-improving systems.
We’re bringing together researchers and engineers working across these layers for five focused technical talks, followed by a speaker panel and audience Q&A.
The first hour is for food, drinks and networking.
Capacity is limited, so please only register if you plan to attend.
Speakers & Talks
Max Shaposhnikov - Research Engineer, Google DeepMind
The 4 Fronts of Agent Capability: Architecture, Post-Training, Harness, and Context
How do you improve an agent when it hits a capability ceiling?
Max will break down four fronts for improvement: model architecture, post-training and RL, harness optimisation, and better context and reusable skills.
Bio: Research Engineer at Google DeepMind working on Gemini, previously at Tessl and Amazon.
Taowen Liu - PhD Researcher, Imperial College London
Agent RL in Practice: From Rollouts to Learning Loops
A practical look at agent RL: how agents generate rollouts, where rewards and feedback come from, how trajectories become training data, and what changes when agents act across long-horizon, tool-using environments.
Bio: Taowen works on LLM training performance, GPU systems and reinforcement learning infrastructure, and is completing a PhD at Imperial College London.
Taowen is also working at Cohere.
Helin Ece Akgül - VP of Engineering, Prolific
Agentic Evals: Trajectory Capture & Human Feedback
How do we evaluate agents beyond a single model response?
Helin will explore trajectory-level signals, human feedback and real-world interaction data, and how these can help identify where agents succeed, fail and improve.
Bio: VP of Engineering at Prolific, previously Director of Engineering at Wise.
Final talk title and description may be updated ahead of the event.
Leon Chlon, PhD - Principal Research Scientist, PhysicsX
KV-Cache Compression and Attention Routing for More Efficient Agents
Leon will share recent work on KV-cache compression and attention routing, including how models can select useful tokens from cached context and use memory more efficiently.
Bio: Principal Research Scientist at PhysicsX and Cambridge-trained AI researcher.
Mike Darlington - Senior Developer Success Engineer, Vercel
fx: Building the Agent Kernel
Inside fx’s kernel-inspired approach to agent harness design - a small portable core that can be embedded, extended or built into larger products.
Bio: Mike works on production agents, coding workflows and agent infrastructure at Vercel, and previously led AI engineering work at Cox Automotive Europe.
Agenda
17:30-18:30 | Doors Open, Food, Drinks & Networking
Check in, grab some food and drinks, and meet researchers, engineers and practitioners working across AI, agents, evaluation and post-training.
18:30-18:40 | Welcome & AAIF London Intro
Opening remarks from Ibrahim Malik, Lead Organiser, AAIF Community London.
A short introduction to AAIF Community London, the wider Agentic AI Foundation, the London organising team, and the evening’s theme around how agents are evaluated and improved after base-model training.
18:40-19:00 | Max Shaposhnikov - Google DeepMind
The 4 Fronts of Agent Capability: Architecture, Post-Training, Harness, and Context
19:00-19:15 | Taowen Liu - Imperial College London
Agent RL in Practice: From Rollouts to Learning Loops
Taowen is also working at Cohere.
19:15-19:30 | Helin Ece Akgül - Prolific
Agentic Evals: Trajectory Capture & Human Feedback
19:30-19:40 | Break
Grab a drink and continue the conversation.
19:40-19:55 | Leon Chlon - PhysicsX
KV-Cache Compression and Attention Routing for More Efficient Agents
19:55-20:10 | Mike Darlington - Vercel
fx: Building the Agent Kernel
20:10-20:40 | Speaker Panel + Audience Q&A
Moderated by Ibrahim Malik, Lead Organiser, AAIF London, and Shaun Smith, MCP / Open Source at Hugging Face.
We’ll bring the speakers together for a hosted discussion on questions including:
What actually makes an agent improve - better models, better post-training, better harnesses, or better context?
Where do current agent evals give us false confidence?
Are we over-investing in larger context windows when better memory and routing might matter more?
Which part of the agent stack is currently the biggest bottleneck to meaningful self-improvement?
We’ll then open the discussion to the audience.
20:40-21:00 | Networking
Continue the conversation with the speakers and attendees.
21:00 onwards | Optional Social
For anyone who wants to continue the evening, we’ll head somewhere nearby.
Location shared on the night.
Who Should Come
This event is for AI/ML engineers, research engineers, researchers, agent developers, applied AI teams and infrastructure engineers working on or interested in post-training, evals, reinforcement learning, agent infrastructure, context, feedback loops and production agentic systems.
If you’re building agents and thinking seriously about how we evaluate, improve and learn from them, the discussions should be relevant.
About Prolific
Prolific provides high-quality human data for AI development and research, helping teams collect reliable human feedback and evaluation data at scale.
Prolific is supporting the event with the venue, food and drinks.
About AAIF
The Agentic AI Foundation (AAIF) is a Linux Foundation initiative supporting an open and interoperable agentic AI ecosystem.
AAIF brings together developers, researchers, open-source maintainers and organisations working across agent protocols, infrastructure, tooling and interoperability.
AAIF Community London runs technical events across agentic AI, machine learning, developer tooling and AI infrastructure.
By attending, you agree to the Linux Foundation Code of Conduct and Privacy Policy.