Cover Image for What It Takes To Build Production-Ready LLM Applications
Cover Image for What It Takes To Build Production-Ready LLM Applications
Avatar for Packt Publishing
Presented by
Packt Publishing

What It Takes To Build Production-Ready LLM Applications

Virtual
Get Tickets
Registration Closed
This event is not currently taking registrations. You may contact the host or subscribe to receive updates.
About Event

​Overview

​Master DSPy to build applications that are accurate, reliable, and production-ready through automated prompt optimization and evaluation

​Most LLM applications start with prompt experiments. The problem is that prompts are hard to test, hard to compare, and hard to improve systematically.

​This hands-on, 3-hour workshop shows you how to move beyond ad hoc prompt engineering and build LLM applications using a more reliable engineering workflow. You’ll learn how to define structured LLM tasks, create measurable baselines, evaluate outputs, identify failure modes, optimize prompts and examples automatically, and track experiments using DSPy and MLflow.

​DSPy is a Python framework for programming, evaluating, and optimizing LLM behavior. Instead of manually rewriting brittle prompt templates, you’ll define signatures, modules, metrics, and optimizers that help you improve LLM application performance in a repeatable way.

​During the workshop, you’ll build a DSPy-powered LLM classifier using a hands-on classification workflow based on the ATIS dataset. You’ll start with a baseline classifier, define the task, prepare evaluation data, measure performance, inspect failures, and then improve the system using few-shot and instruction-level optimization.

​This is not a beginner prompt-engineering session. It is designed for professionals who want to build LLM applications that can be tested, optimized, tracked, explained, and adapted for production workflows.

​Why This Workshop Is Worth Paying For NOW

​Free prompt-engineering tutorials can show you how to write better prompts. This workshop shows you how to build a repeatable engineering workflow for improving LLM applications.

​As teams move from prototypes to production, the hard questions are no longer just “what prompt works?” They are: How do we evaluate quality? How do we compare versions? How do we reduce failures? How do we track changes? How do we make improvements repeatable?

​This workshop gives you a practical framework for answering those questions using DSPy, evaluation metrics, optimizers, and MLflow-based experiment tracking.

​You’ll learn how to:

  • ​Move from manual prompt tweaking to structured LLM application development

  • ​Define LLM tasks using DSPy signatures and modules

  • ​Build a baseline LLM classifier and measure its performance

  • ​Create evaluation examples and task-specific metrics

  • ​Identify failure modes using evaluation results and traces

  • ​Apply few-shot optimization to improve LLM behavior

  • ​Use instruction-level optimization to improve prompts systematically

  • ​Compare baseline and optimized results

  • ​Track experiments and LLM traces with MLflow

  • ​Save, load, and reuse optimized DSPy programs

  • ​Understand where DSPy fits in production LLM application workflows

  • ​Explain LLM reliability, evaluation, and optimization to technical and business stakeholders

​What You’ll Build

​You’ll build and optimize a complete DSPy-powered LLM classification workflow.

​You’ll start by defining a structured LLM task using DSPy signatures and modules. You’ll then create a baseline classifier, prepare evaluation examples, define a metric, and run your first evaluation. After that, you’ll inspect the results, identify failure modes, and apply DSPy optimizers to improve performance using few-shot examples and prompt instruction optimization.

​You’ll also explore how to track experiments and LLM traces with MLflow, compare pre- and post-optimization results, and save optimized DSPy programs so they can be reused in future applications.

​The workflow is based on the ATIS dataset, but the pattern applies to many real-world LLM use cases, including intent classification, ticket routing, customer support triage, document classification, compliance review, and enterprise workflow automation.

​What You'll Get

​By the end of the workshop, you will have:

  • ​Certificate of completion

  • ​Access to full HD event recording

  • ​Reusable blueprint for building more reliable LLM applications

  • ​Build a hands-on DSPy workflow

  • ​Create a baseline LLM classifier

  • ​Design an evaluation set and metric for an LLM task

  • ​Apply few-shot and prompt instruction optimization

  • ​Compare pre- and post-optimization results

  • ​Learn how to save and load optimized DSPy programs

  • ​Explore MLflow for experiment tracking and LLM traces

  • ​How to frame LLM reliability for technical and business stakeholders

​Who Should Attend

​This workshop is ideal for:

  • ​ML engineers and applied AI developers

  • ​Data scientists building LLM workflows

  • ​AI product managers and innovation leads

  • ​Technical founders and CTOs

  • ​Enterprise teams evaluating GenAI use cases

  • ​Consultants and solution architects building AI applications for clients

​About the Speakers

​Serj Smorodinsky and Brett Kennedy are AI engineers and co-authors of a book on building LLM applications with DSPy. Their work focuses on helping teams move beyond brittle static prompts toward adaptive, contract-based LLM applications that can be evaluated, optimized, and improved systematically. They bring practical experience across software development, data science, fraud detection, financial auditing, explainable anomaly detection, DSPy workflows, LLM evaluation, and production-oriented AI engineering.

Avatar for Packt Publishing
Presented by
Packt Publishing