Avatar for THE INFERENCE HUB
Presented by
THE INFERENCE HUB
158 Going

Your R&D Bottleneck Has Moved: How to Get More R&D Output from People and Tokens

Zoom
Registration
Welcome! To join the event, please register below.
About Event

The Agentic Night Shift: Are You Paying Interactive Prices for Agents Nobody Is Interacting With?

Elad Guttel
Co-Founder and VP of R&D

Abstract

Coding is rapidly changing shape. A year ago, the unit of AI work was a prompt, and the core metric was time-to-first-token. Today, the unit is a task: an agent picks up a ticket, a Sentry alert, or a flaky test, works independently for half an hour, and hands back a PR for review.

Stripe merges more than a thousand of these PRs each week with zero human-written code, while METR's time-horizon research shows that the length of tasks AI agents can complete is doubling every few months. Yet almost all of these background workloads still rely on the most expensive tokens on the market: interactive, low-latency frontier models, even when nobody is waiting for the result.

This talk explores that mismatch. We will walk through the inference Pareto curve, from bulk tokens and the "Goldilocks zone" to premium low-latency inference, explain why model providers are introducing Flex, Batch, and Priority tiers, and put concrete numbers on the trade-offs.

What is the value of completing a task in 30 minutes instead of 35 when human review, rather than model inference, is the actual bottleneck?

Elad will also share lessons from large-scale coding-agent workloads, including SWE-bench experiments that achieved comparable pass rates at dramatically lower cost per task. The talk will examine where lower-cost inference falls short, where interactive inference still earns its premium, and how a dedicated "night shift" lane could fundamentally change how engineering organizations purchase compute.

About Elad

Elad Guttel is a Co-Founder and VP of R&D focused on building infrastructure for efficient, large-scale AI inference.


Engineering AI for Production: From Experiments to Autonomous Systems

Renen Avneri
AI Architect, Riverside

Discussion

As AI agents take on increasingly complex work, the challenge is no longer simply making them capable. It is making them reliable, observable, secure, and economically viable in production.

In this discussion, Renen will explore what it takes to move from LLM experiments to production-grade autonomous systems. The conversation will cover agent architecture, observability, control, feedback loops, cost management, and the role of AI within the software development lifecycle.

Drawing on his experience building and operating these systems at scale, Renen will also discuss how organizations can give AI agents greater autonomy without losing visibility or control.

About Renen

Renen Avneri is an AI Architect at Riverside, focused on building production-grade AI systems and autonomous agents. His work spans architecture, observability, security, and cost optimization, with a focus on making AI reliable and practical at organizational scale.

Avatar for THE INFERENCE HUB
Presented by
THE INFERENCE HUB
158 Going