Cover Image for Running open models in production: a live walkthrough of our new inference platform
Cover Image for Running open models in production: a live walkthrough of our new inference platform
Avatar for Together AI Calendar

Running open models in production: a live walkthrough of our new inference platform

YouTube
Registration
Welcome! To join the event, please register below.
About Event

Open-weight models give you real control over quality, performance, and cost. Getting that control into production usually means building a platform team's worth of infrastructure first.

Join us for a live look at the biggest update to Together Dedicated Model Inference since launch. We'll walk through how to bring any model to production in minutes, then roll it out, test it on real traffic, and scale it to your SLOs without rebuilding your deployment along the way.

What we'll cover:

  • Deploying a model or adapter from Hugging Face, S3, or local, with ~4x faster warm starts

  • Safe rollouts with canary, blue-green, and automatic rollback

  • Testing new versions on live traffic with shadow requests and A/B tests

  • SLO-driven autoscaling and multi-region deployment behind one stable endpoint

  • Live Q&A with the team building it

Speakers:

  • Nikitha Suryadevara, Inference Product Manager, Together AI

  • Zain Hasan, Staff AI/ML Engineer, Together AI

Avatar for Together AI Calendar