

Forward Deployed: Evals - Beyond the Vibe Check
Most teams still evaluate their AI the same way: run a few prompts, eyeball the outputs, and ship if it feels right.
That works until it doesn't — until the agent fails silently in production, the hallucination reaches a customer, or nobody can explain why last week's change made things worse.
This fireside gets into what comes after the vibe check: what a production eval actually is and how to build one, where eval datasets come from, using online evals as live guardrails, the LLM-as-judge vs. purpose-built eval models debate, and whether you can grade a multi-step agent's whole trajectory or just its final answer. Plus the practical calls every team faces — build vs. buy, and who owns evals inside a company.
🎤 🎤 🎤 SPEAKERS 🎤 🎤 🎤
LIAM BUSH | Deployed Engineer @ LangChain
LangChain is the default open-source framework for building with LLMs — 118,000+ GitHub stars, $1.25B valuation, powering agents and eval pipelines for thousands of companies — and the team behind LangSmith for evals and observability.LOTTE VERHEYDEN | Developer Relations @ Langfuse
Langfuse is the leading open-source platform for LLM observability and evals — 26M+ SDK installs a month, trusted by 19 of the Fortune 50, and acquired by ClickHouse in early 2026. The open-source, self-hostable tool that became a developer favorite for measuring LLM quality.SOUMYA MOHAN | Head of Product @ Galileo
Galileo is an evaluation platform built to make AI agents reliable — catching hallucinations and failures before production. $68M raised, 834% revenue growth in a year, with research-backed metrics that score AI output automatically, at scale.JULIA ROSE | Staff Product Manager @ Weights & Biases (CoreWeave)
CoreWeave is the AI Hyperscaler powering the world's top AI labs, with 2025 revenue guided near $5B. Its Weights & Biases platform brings rigor to LLM evaluation, with Weave monitoring production AI agents in real time across any cloud.BRADEN HOLSTEGE | VP, Enterprise AI @ Mercor
Mercor is the platform powering the LLMs you use every day — connecting frontier labs with the experts who train and evaluate their models. Revenue went from $75M to $850M in under 8 months. Evaluation is at the core of how frontier models get better.
🏠 HOSTED BY FORWARD DEPLOYED
Forward Deployed is an AI-Native startup building agentic products for PE portfolio companies, Insurers, and other Enterprise teams. We also host fireside chats and podcasts with operators at the forefront of forward deployed engineering.
📅 Tue, Sep 1st, 6-8:30pm PT
📍 Super secret, San Francisco
🎟 Seats are limited
✌️ Hosted at the Founders Cafe, AngelList's founder community and co-working space on the first floor of AngelList HQ. Apply here.