Cover Image for Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMops
Cover Image for Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMops
Avatar for Packt Publishing
Presented by
Packt Publishing
3 Going

Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMops

Virtual
Get Tickets
Welcome! Please choose your desired ticket type:
About Event

Overview

Live, hands-on 3.5-hour masterclass: ship LLM systems that survive production — rigorous evals, RAG, agents, cost & latency control.

Your prompt tweak shipped on Friday. Quality quietly dropped all week — and a customer noticed before you did.

That's how most teams run LLM features today: model choices by gut feel, "evals" that are just vibe checks, retrieval that's never measured, agents that compound their own errors, and cost surprises that arrive as invoices.

This live, hands-on masterclass replaces guesswork with engineering discipline. In 3.5 hours, you'll build a production LLM workflow end to end — versioned prompts, a real evaluation harness, statistically sound model comparisons, evaluated RAG, resilient tool-using agents, and the observability to run it all in production.

You leave with working code: runnable notebooks, templates, and a production-readiness checklist you can apply on Monday morning.

Who this is for

  • Software and ML engineers shipping (or about to ship) LLM-powered features

  • Data scientists moving LLM prototypes from notebook to production

  • Tech leads who need defensible answers to "which model?", "did that change actually help?" and "why is this so expensive?"

  • The hands-on work runs in Python notebooks against real model APIs, so you'll get the most out of it if you're comfortable with Python and JSON. This is not a no-code overview or an AI-trends talk — you'll work with real code, live.

What you will build

  • A versioned prompt pipeline — reusable templates, structured outputs and regression tests, so a prompt edit can never silently degrade quality again

  • A golden dataset and automated eval harness — deterministic checks plus rubric-based LLM-as-judge, so "did the new model help?" gets answered by data, not opinions

  • A statistically rigorous model-comparison workflow — bootstrap confidence intervals and paired significance tests, so you can defend an upgrade with numbers instead of anecdotes

  • An evaluated RAG pipeline — embeddings, vector retrieval, reranking, recall@k and MRR, so you know retrieval works before your users find out it doesn't

  • A tool-using agentic workflow — function calling, validation, guardrails, retries and fallbacks, so your agent fails gracefully instead of hallucinating through errors

  • A production operations layer — tracing, token/cost monitoring, latency budgets and caching, so problems show up in dashboards, not invoices

  • A CLI regression suite you can wire into CI — so quality regressions are caught before deployment, not after

Tools and techniques you will learn

  • Model APIs: OpenAI & Anthropic APIs, structured outputs & schema validation, function calling, streaming, prompt caching

  • Prompt & eval engineering: prompt versioning and templates, deterministic checks, LLM-as-judge, bootstrap confidence intervals, paired significance testing, CLI regression suites

  • RAG: embedding models, vector stores, dense retrieval, reranking, context assembly, retrieval evaluation with recall@k and MRR

  • LLMOps: tracing and observability, token/cost/latency monitoring, caching, retries, fallbacks, graceful degradation

Why this is not another RAG workshop

Plenty of workshops end where production begins — a notebook that works once, on a sunny day. This one covers the whole lifecycle:

  • Choose models on quality, cost and latency — not leaderboard hype

  • Treat prompts as versioned, regression-tested software

  • Replace "it feels better" with paired statistical tests

  • Make RAG architecture decisions you can justify

  • Build agents that recover from errors instead of compounding them

  • Trace, monitor and control cost after you ship — not just before

What you will receive

  • To keep: certificate of completion, full-HD recording, the complete slide deck, lifetime access to all materials

  • To run: runnable notebooks and code repos, golden-dataset and eval-harness templates, the prompt-versioning blueprint with regression-test scaffolding, function-calling and guardrail templates, a CLI regression-suite starter for your CI

  • To decide with: the RAG architecture decision framework, production-readiness checklist, and observability, caching, fallback and cost-control patterns

  • Plus live access to the instructor for your own architecture and implementation questions.

  • Can't make it live? Every ticket includes the full recording and lifetime access to the materials — register anyway and watch on your schedule.

Why this workshop now

LLM capabilities are advancing faster than most teams’ engineering standards. It is increasingly easy to produce an impressive prototype. The difficult part is knowing whether a change actually improved the system—and whether that system will remain reliable when the model, prompt, data, traffic, or API behaviour changes.

Without a disciplined production workflow:

  • Prompt changes introduce silent regressions.

  • Model upgrades become subjective “looks better” decisions.

  • Retrieval systems return plausible but poorly grounded answers.

  • Agents compound errors across multiple tool calls.

  • Cost and latency issues only appear after usage grows.

  • Teams cannot explain why quality changed between releases.

  • Production incidents are difficult to reproduce or diagnose.

The durable advantage is therefore not expertise in one vector database, agent framework, or model provider. It is the ability to measure quality, compare alternatives, trace failures, control operational trade-offs, and improve the system safely.

This masterclass gives engineers that complete production discipline—from the first architecture decision to the evaluation suite and observability needed after deployment.

Learn from someone who's done this before: About the speaker

Bruno Gonçalves is the founder of Data For Science. As an author, speaker, corporate trainer and consultant specialising in Generative AI, machine learning and production LLM systems, he has trained hundreds of engineers at Fortune 500 companies. Previously a Data Science Fellow at NYU's Centre for Data Science and tenured faculty at Aix-Marseille Université, he holds a PhD in the Physics of Complex Systems. The statistical rigour in this masterclass isn't a buzzword.

Reserve your seat

Saturday 12 September · 9:30 AM–1:00 PM EDT (2:30 PM UK · 3:30 PM CEST · 7:00 PM IST) · Live online

Full refunds up to 7 days before the event — if it turns out not to be for you, you get your money back.

Get your ticket and come build LLM systems you can actually trust.

Avatar for Packt Publishing
Presented by
Packt Publishing
3 Going