

What It Takes To Build Production-Ready LLM Applications
Overview
Master DSPy to build applications that are accurate, reliable, and production-ready through automated prompt optimization and evaluation
Most LLM applications start with prompt experiments. The problem is that prompts are hard to test, hard to compare, and hard to improve systematically.
This hands-on, 3-hour workshop shows you how to move beyond ad hoc prompt engineering and build LLM applications using a more reliable engineering workflow. You’ll learn how to define structured LLM tasks, create measurable baselines, evaluate outputs, identify failure modes, optimize prompts and examples automatically, and track experiments using DSPy and MLflow.
DSPy is a Python framework for programming, evaluating, and optimizing LLM behavior. Instead of manually rewriting brittle prompt templates, you’ll define signatures, modules, metrics, and optimizers that help you improve LLM application performance in a repeatable way.
During the workshop, you’ll build a DSPy-powered LLM classifier using a hands-on classification workflow based on the ATIS dataset. You’ll start with a baseline classifier, define the task, prepare evaluation data, measure performance, inspect failures, and then improve the system using few-shot and instruction-level optimization.
This is not a beginner prompt-engineering session. It is designed for professionals who want to build LLM applications that can be tested, optimized, tracked, explained, and adapted for production workflows.
Why This Workshop Is Worth Paying For NOW
Free prompt-engineering tutorials can show you how to write better prompts. This workshop shows you how to build a repeatable engineering workflow for improving LLM applications.
As teams move from prototypes to production, the hard questions are no longer just “what prompt works?” They are: How do we evaluate quality? How do we compare versions? How do we reduce failures? How do we track changes? How do we make improvements repeatable?
This workshop gives you a practical framework for answering those questions using DSPy, evaluation metrics, optimizers, and MLflow-based experiment tracking.
You’ll learn how to:
Move from manual prompt tweaking to structured LLM application development
Define LLM tasks using DSPy signatures and modules
Build a baseline LLM classifier and measure its performance
Create evaluation examples and task-specific metrics
Identify failure modes using evaluation results and traces
Apply few-shot optimization to improve LLM behavior
Use instruction-level optimization to improve prompts systematically
Compare baseline and optimized results
Track experiments and LLM traces with MLflow
Save, load, and reuse optimized DSPy programs
Understand where DSPy fits in production LLM application workflows
Explain LLM reliability, evaluation, and optimization to technical and business stakeholders
What You’ll Build
You’ll build and optimize a complete DSPy-powered LLM classification workflow.
You’ll start by defining a structured LLM task using DSPy signatures and modules. You’ll then create a baseline classifier, prepare evaluation examples, define a metric, and run your first evaluation. After that, you’ll inspect the results, identify failure modes, and apply DSPy optimizers to improve performance using few-shot examples and prompt instruction optimization.
You’ll also explore how to track experiments and LLM traces with MLflow, compare pre- and post-optimization results, and save optimized DSPy programs so they can be reused in future applications.
The workflow is based on the ATIS dataset, but the pattern applies to many real-world LLM use cases, including intent classification, ticket routing, customer support triage, document classification, compliance review, and enterprise workflow automation.
What You'll Get
By the end of the workshop, you will have:
Certificate of completion
Access to full HD event recording
Reusable blueprint for building more reliable LLM applications
Build a hands-on DSPy workflow
Create a baseline LLM classifier
Design an evaluation set and metric for an LLM task
Apply few-shot and prompt instruction optimization
Compare pre- and post-optimization results
Learn how to save and load optimized DSPy programs
Explore MLflow for experiment tracking and LLM traces
How to frame LLM reliability for technical and business stakeholders
Who Should Attend
This workshop is ideal for:
ML engineers and applied AI developers
Data scientists building LLM workflows
AI product managers and innovation leads
Technical founders and CTOs
Enterprise teams evaluating GenAI use cases
Consultants and solution architects building AI applications for clients
About the Speakers
Serj Smorodinsky and Brett Kennedy are AI engineers and co-authors of a book on building LLM applications with DSPy. Their work focuses on helping teams move beyond brittle static prompts toward adaptive, contract-based LLM applications that can be evaluated, optimized, and improved systematically. They bring practical experience across software development, data science, fraud detection, financial auditing, explainable anomaly detection, DSPy workflows, LLM evaluation, and production-oriented AI engineering.