

Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMops
Overview
Live, hands-on 3.5-hour masterclass: ship LLM systems that survive production — rigorous evals, RAG, agents, cost & latency control.
Your prompt tweak shipped on Friday. Quality quietly dropped all week — and a customer noticed before you did.
That's how most teams run LLM features today: model choices by gut feel, "evals" that are just vibe checks, retrieval that's never measured, agents that compound their own errors, and cost surprises that arrive as invoices.
This live, hands-on masterclass replaces guesswork with engineering discipline. In 3.5 hours, you'll build a production LLM workflow end to end — versioned prompts, a real evaluation harness, statistically sound model comparisons, evaluated RAG, resilient tool-using agents, and the observability to run it all in production.
You leave with working code: runnable notebooks, templates, and a production-readiness checklist you can apply on Monday morning.
Who this is for
Software and ML engineers shipping (or about to ship) LLM-powered features
Data scientists moving LLM prototypes from notebook to production
Tech leads who need defensible answers to "which model?", "did that change actually help?" and "why is this so expensive?"
The hands-on work runs in Python notebooks against real model APIs, so you'll get the most out of it if you're comfortable with Python and JSON. This is not a no-code overview or an AI-trends talk — you'll work with real code, live.
What you will build
A versioned prompt pipeline — reusable templates, structured outputs and regression tests, so a prompt edit can never silently degrade quality again
A golden dataset and automated eval harness — deterministic checks plus rubric-based LLM-as-judge, so "did the new model help?" gets answered by data, not opinions
A statistically rigorous model-comparison workflow — bootstrap confidence intervals and paired significance tests, so you can defend an upgrade with numbers instead of anecdotes
An evaluated RAG pipeline — embeddings, vector retrieval, reranking, recall@k and MRR, so you know retrieval works before your users find out it doesn't
A tool-using agentic workflow — function calling, validation, guardrails, retries and fallbacks, so your agent fails gracefully instead of hallucinating through errors
A production operations layer — tracing, token/cost monitoring, latency budgets and caching, so problems show up in dashboards, not invoices
A CLI regression suite you can wire into CI — so quality regressions are caught before deployment, not after
Tools and techniques you will learn
Model APIs: OpenAI & Anthropic APIs, structured outputs & schema validation, function calling, streaming, prompt caching
Prompt & eval engineering: prompt versioning and templates, deterministic checks, LLM-as-judge, bootstrap confidence intervals, paired significance testing, CLI regression suites
RAG: embedding models, vector stores, dense retrieval, reranking, context assembly, retrieval evaluation with recall@k and MRR
LLMOps: tracing and observability, token/cost/latency monitoring, caching, retries, fallbacks, graceful degradation
Why this is not another RAG workshop
Plenty of workshops end where production begins — a notebook that works once, on a sunny day. This one covers the whole lifecycle:
Choose models on quality, cost and latency — not leaderboard hype
Treat prompts as versioned, regression-tested software
Replace "it feels better" with paired statistical tests
Make RAG architecture decisions you can justify
Build agents that recover from errors instead of compounding them
Trace, monitor and control cost after you ship — not just before
What you will receive
To keep: certificate of completion, full-HD recording, the complete slide deck, lifetime access to all materials
To run: runnable notebooks and code repos, golden-dataset and eval-harness templates, the prompt-versioning blueprint with regression-test scaffolding, function-calling and guardrail templates, a CLI regression-suite starter for your CI
To decide with: the RAG architecture decision framework, production-readiness checklist, and observability, caching, fallback and cost-control patterns
Plus live access to the instructor for your own architecture and implementation questions.
Can't make it live? Every ticket includes the full recording and lifetime access to the materials — register anyway and watch on your schedule.
Why this workshop now
LLM capabilities are advancing faster than most teams’ engineering standards. It is increasingly easy to produce an impressive prototype. The difficult part is knowing whether a change actually improved the system—and whether that system will remain reliable when the model, prompt, data, traffic, or API behaviour changes.
Without a disciplined production workflow:
Prompt changes introduce silent regressions.
Model upgrades become subjective “looks better” decisions.
Retrieval systems return plausible but poorly grounded answers.
Agents compound errors across multiple tool calls.
Cost and latency issues only appear after usage grows.
Teams cannot explain why quality changed between releases.
Production incidents are difficult to reproduce or diagnose.
The durable advantage is therefore not expertise in one vector database, agent framework, or model provider. It is the ability to measure quality, compare alternatives, trace failures, control operational trade-offs, and improve the system safely.
This masterclass gives engineers that complete production discipline—from the first architecture decision to the evaluation suite and observability needed after deployment.
Learn from someone who's done this before: About the speaker
Bruno Gonçalves is the founder of Data For Science. As an author, speaker, corporate trainer and consultant specialising in Generative AI, machine learning and production LLM systems, he has trained hundreds of engineers at Fortune 500 companies. Previously a Data Science Fellow at NYU's Centre for Data Science and tenured faculty at Aix-Marseille Université, he holds a PhD in the Physics of Complex Systems. The statistical rigour in this masterclass isn't a buzzword.
Reserve your seat
Saturday 12 September · 9:30 AM–1:00 PM EDT (2:30 PM UK · 3:30 PM CEST · 7:00 PM IST) · Live online
Full refunds up to 7 days before the event — if it turns out not to be for you, you get your money back.
Get your ticket and come build LLM systems you can actually trust.