

Running open models in production: a live walkthrough of our new inference platform
Open-weight models give you real control over quality, performance, and cost. Getting that control into production usually means building a platform team's worth of infrastructure first.
Join us for a live look at the biggest update to Together Dedicated Model Inference since launch. We'll walk through how to bring any model to production in minutes, then roll it out, test it on real traffic, and scale it to your SLOs without rebuilding your deployment along the way.
What we'll cover:
Deploying a model or adapter from Hugging Face, S3, or local, with ~4x faster warm starts
Safe rollouts with canary, blue-green, and automatic rollback
Testing new versions on live traffic with shadow requests and A/B tests
SLO-driven autoscaling and multi-region deployment behind one stable endpoint
Live Q&A with the team building it
Speakers:
Nikitha Suryadevara, Inference Product Manager, Together AI
Zain Hasan, Staff AI/ML Engineer, Together AI