

How to deploy a multi-model architecture
For the past few years, companies have turned to large frontier models to build their AI applications, as they could address most use cases and were the only viable options at scale. But this world is changing fast, and today, the teams shipping the best AI products are increasingly turning to smaller, specialized models because they beat the generalists on core tasks, are often much faster, and can be deployed at a fraction of the cost.
Some common examples we think of are voice AI that can respond in real time, fast retrieval in structured data documents, or predictions off tables businesses run on.
Join Baseten and leaders from Gradium, Synthefy, and Sid.ai for drinks, snacks, fun demos, and a candid conversation about why the industry is shifting toward a multi-model world and how Baseten can help power it.
We are bringing together leaders at three labs building for this new world:
Constance Grisoni: Chief Growth Officer of Gradium - A real-time text-to-speech, speech-to-text and voice cloning at ultra-low latency.
Somi Agarwal: Co-Founder of Synthefy - A foundation model for structured data that enables predictions and forecasts straight from tables.
Max Rumpf: Co-Founder of Sid.ai - An agent search model specialized in retrieval that outperforms frontier models on recall while running an order of magnitude faster and cheaper.
The panel will be moderated by Bola Malek, Head of Baseten for Model Labs. They will go into what their models do that a general-purpose model can’t, where the specialist vs. generalist line falls, and what it takes to run a fleet of models in production efficiently and at scale.
This event is hosted by Baseten — the training and inference platform behind some of the most demanding AI workloads in production like Abridge, Cursor, Notion, and OpenEvidence. Most bring their own models. We make them fast.