Gavel & Gown: LLM Judge
βYour agent's output "looks good." Can you prove it? π¨
βThat's the gap this session closes. Installment 3 of the series is a laptop-open build for anyone shipping agent output they can't actually verify.
βHere's the problem: "looks good" is a vibe, not a QA process. So bad answers slip through, good ones get thrown out, and nobody can say why. The fix isn't a better prompt. It's a judge.
βSo we build one. An unbiased LLM-as-Judge that turns "seems fine" into a measurable keep-or-revert verdict.
β(Your Judge will be judged by Judy the Judge. Kidding. Mostly.)
βYou'll walk out knowing how to build:
βA deterministic process so the same input always gets the same ruling
βEvals that score outputs against a rubric instead of a gut feeling
βLoops that feed the verdict back in so the system sharpens itself
βThis is the QA layer that separates agents you can trust from agents you just hope about.
βBring your laptop. Leave with a judge.
βBest for: Founders, engineers, AI operators, and anyone building agents that need to be right, not just plausible.
βπ Antler VC β 800 Brazos St #340, Austin, TX
π₯οΈ Austin in-person energy + recordings (free for members)
ποΈ Approval required β request your spot early
βRequest to Join π