

Presented by
Warp
Mostly livestreamed, sometimes in person events by the Warp team. Learn more about what we're building, how we're using it, and who we are.
Hosted By
3 Going
How to benchmark the best models for your team's cost-per-PR
Registration
About Event
You can reduce your cost-per-PR significantly by optimizing model selections. But how do you decide which models to use across different tasks? Public benchmarks like SWEBench? Manual trial-and-error? The Twitter flavor of the day?
From our testing, the best tool isn’t public benchmarks or recommendations; it's testing on your team's own coding tasks.
In this session, I’ll show how to build a benchmark by replaying your team’s past agent runs using Factory Benchmarks. We'll walk through how to find the right sample data, how to score benchmark runs, and how to understand a benchmark report to decide the best model to reduce your team’s cost-per-PR.
We used this approach ourselves to drive down our costs from $80 per PR to $30.
Presented by
Warp
Mostly livestreamed, sometimes in person events by the Warp team. Learn more about what we're building, how we're using it, and who we are.
Hosted By
3 Going