

Reliable AI in Production: Evals, Observability and Langfuse
Date: Monday, September 7th, 2026
Time: 18:00 to 20:00
Location: Tikal Offices, Kremnitski 6, Tel Aviv
*The talks will be in Hebrew
Getting an LLM application to work is one thing. Understanding how it behaves in production, identifying why it failed, and preventing the same failure from happening again is a different challenge.
As AI applications evolve from simple prompts into agentic workflows, teams need more than logs and occasional manual testing. They need evaluation systems that define what good behavior looks like, alongside observability that reveals what happened across every prompt, tool call, and model interaction.
This meetup will focus on these two connected layers of reliable AI systems.
Keren Finkelstein will present a practical methodology for building an evaluation system, from manual exploration and clear pass or fail criteria to automated graders, test datasets, and CI/CD integration.
Chaim Turkel will then show how Langfuse can provide visibility into agent execution, prompts, tool usage, latency, token consumption, and cost, while supporting continuous evaluation in development and production.
On the Agenda
18:00 to 18:30
Welcome Drinks & Networking
18:30 to 19:15
AI Evals Best Practices: Keeping AI Behavior Reliable in Production// Keren Finkelstein, Backend Tech Lead and Group Lead, Tikal
Shipping an LLM application is often the easy part. Keeping its behavior reliable, safe, and trustworthy as the application and its data evolve is where many teams struggle.
In this session, Keren will walk through a practical methodology for building an evaluation system. She will cover manual exploration, defining clear pass or fail criteria, and combining structural graders with semantic graders powered by LLM as a judge.
The session will follow the full evaluation lifecycle, from building a representative test dataset to integrating evaluations into CI/CD. Keren will also discuss the patterns that keep evaluation systems useful over time, and how production failures can become part of a continuous feedback loop that improves the system.
19:15 to 20:00
Observability for Agentic AI with Langfuse// Chaim Turkel, Agentic AI and Data Architect & Group Lead, Tikal
Building an LLM application is only the beginning. Understanding what it actually does in production is what makes it possible to debug, evaluate, and improve.
In this session, Chaim will demonstrate how Langfuse provides end to end observability for LLM and agentic applications. He will show how to inspect the exact prompts sent to models, trace agent execution, understand which tools were invoked and why, and monitor latency, token usage, and cost across every step of an interaction.
The session will also cover evaluation datasets, automated prompt testing, and scoring pipelines that measure application quality during development and in production. The goal is to show how tracing, evaluation, and analytics work together to make AI applications more observable, measurable, and reliable.
Please Note
Registering through this page is an application for approval and will be reviewed by the team before tickets are confirmed.
About Tikal
Tikal is a hands-on tech consultancy living and breathing the Agentic world.
For 25+ years we've scaled engineering organizations, today we help them pioneer Agentic Engineering, where developers and AI agents build together. AI is dissolving the old silos between disciplines, and that's exactly where we're strongest: deep, cross-domain expertise in AI/ML, Backend, Data, DevOps, Fullstack, and
Web, integrating in your teams and delivering value from day one.
As the founders of the Israeli Tech Radar, we drive industry impact by sharing insights and guiding engineering leaders and teams across the tech industry.