

Round table: You can measure what AI costs. Can you measure what it returned?
Generative AI is live in production across most large enterprises. Spend is visible, usage dashboards are full, and boards have started asking the harder question: what did it actually return?
The people accountable can report message volume and inference cost, but not what actually came back: time saved and revenue influenced. And there is a second question underneath the first. Improving that return is only half about the model. The other half is how skilled people are at using it, which almost nobody measures. The root cause is a blind spot the whole market shares. Every tool instruments the model (latency, tokens, traces), and almost none can see the human on the other end: what they were trying to do, whether they got it, and where the experience quietly failed them. This table is for the leaders now on the hook for that answer, ahead of the next budget review.
What we'll discuss:
- Why usage and cost metrics stop being enough the moment AI moves from pilot to scaled rollout
- What a credible answer to "what did it return" actually rests on: time saved on real tasks, and the revenue signals sitting in customer-facing conversations
- Why improving that return is only half about the model, and how measuring user proficiency, which almost nobody does, is the other half
- The gap between observing the system and understanding the user, and why that gap is where adoption plateaus, churn signals, and unrealized value hide
- Why the less than 2% of users who leave a rating are a poor basis for decisions about a system handling millions of interactions
- What an ROI number that survives a board meeting looks like, and what it takes to produce one
- Doing this inside regulated environments where prompt data cannot leave the perimeter, such as banks, insurers, and telcos
Hosted by Nebuly