

Presented by
Warp
Mostly livestreamed, sometimes in person events by the Warp team. Learn more about what we're building, how we're using it, and who we are.
Hosted By
42 Going
Reducing your team's cost-per-PR with custom benchmarks
Registration
About Event
You can reduce your cost-per-PR significantly by optimizing model selections. But how do you decide which models to use across different tasks? Public benchmarks like SWEBench? Manual trial-and-error? The Twitter flavor of the day?
From our testing, the best tool isn’t public benchmarks or recommendations; it's testing on your team's own coding tasks.
In this session, I’ll show how to build a benchmark by replaying your team’s past agent runs using Factory Benchmarks. We'll walk through how to find the right sample data, how to score benchmark runs, and how to understand a benchmark report to decide the best model to reduce your team’s cost-per-PR.
We used this approach ourselves to drive down our costs from $80 per PR to $30.
Presented by
Warp
Mostly livestreamed, sometimes in person events by the Warp team. Learn more about what we're building, how we're using it, and who we are.
Hosted By
42 Going