

Deterministic AI on CPUs: Build low-latency LLM workflows
Build deterministic, low-latency AI workflows on CPUs using encoders, conformal prediction, LangGraph and selective LLM escalation.
This build lab starts where standard agent tutorials break down: when an LLM orchestrator suffers from context drift, schema failures, and compounding latency in production. Using an enterprise triage and claims-processing case study, you will replace generative routing loops with bidirectional encoders (LFM2.5 and ModernBERT), adapt them to taxonomies using few-shot contrastive learning (SetFit), calibrate uncertainty bounds with conformal prediction, orchestrate idempotent execution in LangGraph, and validate the pipeline using Wilson score intervals and clustered standard errors.
Instead of another prompt-driven agent demo, this workshop equips you with a reusable deterministic-probabilistic architecture that runs locally on standard CPUs, eliminates JSON parsing failures at the routing layer, and restricts LLM API calls to verified edge cases.
What tools and techniques you will learn
Python, NumPy, and Jupyter or Google Colab
Liquid AI LFM2.5-Encoder and ModernBERT for long-context CPU inference
Zero-shot prompt routing and vector scoring heads
Few-shot contrastive fine-tuning with SetFit
Split conformal prediction for distribution-free uncertainty calibration
Deterministic state machine orchestration and idempotency in LangGraph
Targeted synthetic hard-negative generation for router hardening
Statistical benchmarking with Wilson score intervals and clustered standard errors
Post-training quantization for CPU throughput optimization
What you will get
Certificate of completion
Full HD recording of the session
Speaker PPT slide deck
Complete runnable notebooks for the enterprise triage and execution engine built in the session.
Prerequisites
You should have intermediate Python skills and familiarity with basic LLM workflows or transformer concepts.
Bring a laptop with a local Python environment or a free Google Colab account.
An OpenRouter account (used for offline synthetic data generation and fallback escalation paths).
About the speaker:
Ben Auffarth is an AI consultant, author, and full-stack data scientist with over 15 years of experience spanning machine learning, generative AI, and large-scale data systems. He holds a PhD in computational and cognitive neuroscience and has worked on systems ranging from brain simulations on IBM supercomputers to real-time ML decision engines processing more than 100,000 transactions daily. Ben is also a best-selling author. His work focuses on practical AI architectures, LLMs, LangGraph, evaluation, and high-performance decision systems, bridging rigorous research with real-world engineering.