Cover Image for How to benchmark the best models for your team's cost-per-PR
Cover Image for How to benchmark the best models for your team's cost-per-PR
Avatar for Warp
Presented by
Warp
Mostly livestreamed, sometimes in person events by the Warp team. Learn more about what we're building, how we're using it, and who we are.
Hosted By
3 Going

How to benchmark the best models for your team's cost-per-PR

Zoom
Registration
Welcome! To join the event, please register below.
About Event

You can reduce your cost-per-PR significantly by optimizing model selections. But how do you decide which models to use across different tasks? Public benchmarks like SWEBench? Manual trial-and-error? The Twitter flavor of the day?

From our testing, the best tool isn’t public benchmarks or recommendations; it's testing on your team's own coding tasks.

In this session, I’ll show how to build a benchmark by replaying your team’s past agent runs using Factory Benchmarks. ​We'll walk through how to find the right sample data, how to score benchmark runs, and how to understand a benchmark report to decide the best model to reduce your team’s cost-per-PR.

We used this approach ourselves to drive down our costs from $80 per PR to $30.

Avatar for Warp
Presented by
Warp
Mostly livestreamed, sometimes in person events by the Warp team. Learn more about what we're building, how we're using it, and who we are.
Hosted By
3 Going