Energy Optimization of GPUs through Self-Improving Agents
Most inference clusters burn far more energy than the work requires. This evening covers the architecture behind a different approach: self-improving agents that keep tuning the serving stack, batching, tensor parallelism, prefix caching, expert routing, until 8 GPUs deliver the throughput you would normally provision 16 for, at roughly half the energy.
Hamza Farooq and the Traversaal team have been building this in the open at energy.traversaal.ai, an inference API priced at a flat $10 per kilowatt-hour instead of per token. Billing on energy changes the incentives: every watt the system saves makes the model genuinely cheaper to call, not just cheaper on the rate card. Their current numbers: 88% under token rates, 0.22s median time to first token.
We will go deep on the underlying architecture: how the agents measure energy per request, what they change between runs, and where the 2x headroom in today's GPU deployments actually comes from.
Who should come: infra and ML engineers, agent builders, researchers, and anyone paying real GPU bills.
Schedule
7:00 PM. Doors open, food and drinks
7:30 PM. Talk: energy-priced inference and the self-improving serving stack
8:15 PM. Open discussion and Q&A
9:00 PM. Wrap
Hosted by Hamza Farooq (Traversaal) and Julius Ritter at AGI House SF.