

What Four Months in Production Did to Maven Clinic's Healthcare AI Agent with William Horton
Four months ago, Maven Assistant had just reached its first external users. AI Engineer William Horton joined Vanishing Gradients after spending the day reading their conversations and discovering what the team had failed to anticipate.
At the time, Maven had released the agent to 20% of users. One capability was still off limits: answering questions about healthcare benefits. The system could not yet clear the reliability bar for decisions that might cost someone tens of thousands of dollars.
Now Maven Assistant is available to 100% of users, and benefits answering has shipped.
William is coming back to compare the agent they launched with the system they have today. We will get into what real users changed, which failures became evals, how the team decided benefits answering was ready, and what a wave of new models did to the architecture.
What We'll Cover
What it took to move Maven Assistant from 20% of users to 100%, and what adoption looked like along the way
How the team decided benefits answering was finally safe enough to ship after withholding it at launch
Which real conversations exposed missing capabilities, broken assumptions, or evaluation gaps
How production failures become regression cases, and which failure changed the system most
What the team expected people to ask for, what they barely used, and which apparently boring requests became essential
How four months of production data changed Maven's eval suite, LLM judges, and synthetic-user testing
How the team evaluates model upgrades without quietly breaking routing, tool use, guardrails, latency, or cost
Which prompts, agents, guardrails, or architectural choices newer models allowed them to simplify or remove
What William would rebuild differently after watching the system operate in the real world
Register to join live or get the recording afterwards.