

Inference Engineering with Modal’s Charles Frye
'Gradient Descending' is a series of roundtables on AI: curated deep dives with technical builders. Past roundtables discussions covered RL Environments, Evals, Fine-Tuning, Agent Frameworks, Open Source Models, and Knowledge Graphs with companies like OpenAI, Anthropic, Cursor, Scale AI and many more.
We’re joined by Charles Frye from Modal for a deep dive into inference engineering.
We’ll get into:
The serving stack optimised to serve models, especially open source models
Achieving Pareto optimal performance through inference optimisation - throughput, latency
How to think about scaling inference without building a huge infra team
As always, expect a deep technical discussion with fellow builders.
Agenda:
Arrivals: 08:30–09:00
Roundtable discussion: 09:00–10:00
See past roundtables here:
https://www.akashbajwa.co/t/ai-roundtables
I agree that photographs and/or video may be taken during the event and used by the hosts on social media and in future promotional materials.