

Harness Engineering: Build the Loop Around Agents You Can Trust
The same model gives one team a flaky demo and another team a dependable agent. The difference is never the model. It's the harness: the system around the model that feeds it context, checks its output, blocks the catastrophic action, and turns every failure into a fix.
Model quality is no longer your differentiator, because everyone rents the same models. Harness quality is, because you build it. This session is a live build of that harness, layer by layer.
We start with a naked model call running a real agent task, and watch it fail: a wrong tool call, confident nonsense, a jailbreak. Then we wrap it:
Add tracing, and the loop becomes visible: what the agent saw, what it called, what it decided, what it cost
Add an eval gate, and the bad output gets caught before it ships
Add a guardrail, and the jailbreak gets blocked mid-stream
Close the loop, and the logged failure becomes a test case that makes the agent measurably better
Same model at the end. Dependable agent. Reliability engineered, not prompted.
Who it's for AI engineers shipping agents, especially if your lived experience is "the model is fine, my system is flaky." AI PMs who own agent quality will get the mental model even if they skip the code.
What you leave with The harness mental model, the runnable open-source repo of the exact harness built live, and a Monday-morning starting order: trace first, then gates, then guardrails, then the loop.
Bring A laptop if you want to run the repo with us. Star github.com/future-agi/future-agi and create a free account at app.futureagi.com before you arrive so you can follow from minute one.
Location: TBA
Free to attend.
RSVPs are approved, so tell us what you're building when you register.