Cover Image for ModelOps Night: Evals, Routing & Benchmarks
Cover Image for ModelOps Night: Evals, Routing & Benchmarks
20 Going

ModelOps Night: Evals, Routing & Benchmarks

Hosted by Renzo, Haritha Nair & Sahar Mor (Bond AI)
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

A technical deep dive on evaluating AI products, agents, and workflows against real-world tasks.

As models get better, the hard question is no longer just “which model is best?” It’s whether your product, agent, or workflow actually performs well on the tasks your users care about.

Public benchmarks rarely capture your edge cases. Real production systems have messy inputs, multi-step workflows, tool calls, latency constraints, cost tradeoffs, and failure modes that only show up in practice.

This event brings together teams building the infrastructure for evaluating AI systems in the real world.

We’ll hear from Composio on building evals for real-world workflows, GMI Cloud on routing models to optimize performance and spend, Handshake on creating high-quality benchmarks that define the frontier, and Oqoqo on how to evaluate and improve agent experience by building your own benchmark.

If you’re building tools for agents, custom agents, or internal workflows, this is a room for practical conversations about how to measure what actually works.

Program

5:30 Doors open · food & drinks

6:00 Composio: building evals for real-world workflows (Jayesh Sharma, AI Engineer at Composio)

6:15 GMI Cloud: routing models to optimize performance and spend (Jay Shu, Product Manager, GMI Cloud)

6:30 Handshake: Building benchmarks for economically-valuable agent work (Jonas Mueller, Director of AI Research, Handshake AI)

6:45 Oqoqo: evaluating and improving agent experience by building your own benchmark (Haritha Nair, CTO, Oqoqo)

7:00 Networking

Who should come

Founders, product engineers, AI engineers, and teams building agents/agent-facing products, internal workflows, or systems that need to be tested against real user tasks.

Location
Werqwise
465 California St 7th floor, San Francisco, CA 94104, USA
20 Going