Cover Image for AI Token cost reduction: Tips and Tricks
Cover Image for AI Token cost reduction: Tips and Tricks
Avatar for antelligent
Presented by
antelligent
We are on a mission to make AI sustainable, private, and deployable everywhere.
63 Going

AI Token cost reduction: Tips and Tricks

Virtual
Registration
Welcome! To join the event, please register below.
About Event

Everyone benchmarks their agent's accuracy. Almost nobody looks at where the tokens actually go.

Here is the uncomfortable answer: in a typical agent run, roughly 94% of what you pay for is input, and most of that is the same context resent over and over. Every clever trick for shortening your model's replies is fighting over the last 2% of the bill.

This meetup goes after the other 98%. In ninety minutes you will learn to read an agent's cost curve, find the three places the money leaks, and fix them live on a real loop, with real numbers on screen.

Bring a laptop and leave with a smaller bill.

Hosted by the team at Fernfly and EverydaySeries, both antelligent AI products. This is the second in our series, the first session built a working agent for $0. This one keeps it that way as it grows.

What you will walk away with

  • A mental model for where agent tokens actually go: prefix, tool schemas, history, tool results, and which one dominates at your number of steps

  • The one-line formula for when your cost curve turns quadratic, and what to do on each side of it

  • Prompt caching done properly: why most teams get a fraction of the 90% discount on offer, and how to fix the cache hygiene that is quietly busting it

  • How to stop 75 tool definitions eating your context window before the agent reads its first message

  • An honest read on which levers are oversold: semantic caching claims of 90% hit rates versus the 20-45% teams actually see in production

  • A working before-and-after on your own agent, measured, not estimated

Agenda

  • Welcome

  • Where the money goes: reading an agent's real cost curve (15 min)

  • The big three: caching, tool schemas, tool result budgets (15 min)

  • When a tiny model does the job: routing steps to Fernfly (10 min)

  • Build: instrument a live agent loop and cut it, together (15 min)

  • Q&A and open floor (20 min)

  • Wrap-up

Who is this for

  • Developers shipping agents who have just met their first surprising invoice

  • Founders and product people trying to make agent unit economics pencil out

  • Anyone running a coding agent, research agent, or long-running workflow and wondering why a short answer costs so much

  • Teams in regulated or cost-constrained settings deciding what to keep in-house

No AI background needed. If you have ever called an LLM API you will follow it fine. Basic comfort with Python or JavaScript helps the build, but you can watch and copy.

Before you arrive (optional, for the hands-on)

  • A laptop

  • A free account at fernfly.com and everydayseries.com

  • An agent, script, or workflow of your own that calls an LLM, even a rough one. We will measure whatever you bring. If you have nothing, use ours.

  • If you can, one month of API usage numbers. Optimising without measuring first is guessing.

One note on timing

Introductory pricing on frontier models by Claude is reported to end 31 August. If that holds, some teams are nine days from a step change in their cost base. Good week to know where your tokens go.

About the hosts

Fernfly turns your API into a tiny model that just calls the right function. It is free to run, private by default, and works everywhere, from a browser to a smartwatch. It is built by antelligent (Advanced Nonlinear Technologies), a team building small, specialist AI that businesses can own.

The same team also builds EverydaySeries, a hosted platform for putting agents to work on everyday tasks while consuming reduced tokens and which is how we learned most of what is in this session the expensive way.

See you Saturday. Bring your invoice.

Avatar for antelligent
Presented by
antelligent
We are on a mission to make AI sustainable, private, and deployable everywhere.
63 Going