

Agent Mode: ON
Meetup Itinerary
17:30 - 18:00 Gathering end registration
18:00 - 18:15 Opening Remarks - Alice Fridberg, Data Science Team Lead, Arpeely
18:15 - 18:45 Not Just Chatbots / Eva Mishor
18:45 - 18:55 Break
18:55 - 19:25 Multi-Agent Systems in the Era of LLMs: Protocols, Patterns, and Agent Architectures / Yaara Cohen
19:25 - 19:55 Evaluating Your AI Agent: How Do You Properly Measure Performance? Shirli Di-Castro Shashua
Talks Abstract:
Not Just Chatbots
Eva Mishor | Founder and CTO at Stealth
AI agents go beyond chatbots; they reason, take actions, and use tools to achieve goals.
This talk breaks down what agents actually are, how they work, and what’s needed to build one.
We’ll look at core architectures, examples, and current frameworks.
Multi-Agent Systems in the Era of LLMs: Protocols, Patterns, and Agent Architectures
Yaara Cohen | R&D Lead
This lecture introduces the foundations and modern developments in multi-agent systems (MAS), with special focus on LLM-based agents and emerging interoperability protocols. It will survey MAS design patterns (e.g. coordinator, blackboard, auction/contract net) and interaction protocols. We will overview and emphasize the connection between MAS and recent standards in generative AI: the Model Context Protocol (MCP), A2A (Agent-to-Agent), and ACP (Agent Communication Protocol), and discuss the challenges these systems arise.
Evaluating Your AI Agent: How Do You Properly Measure Performance?
Shirli Di-Castro Shashua | Senior AI Scientist, Intuit
AI agents are becoming the next big thing. But deploying an agent without truly understanding its performance, limits, and potential failure points is a high-stakes gamble. How do you ensure your agent is not just functional, but genuinely reliable, robust, and safe?
This talk explores the practical challenges of evaluating AI agents effectively. We'll discover how to define meaningful success metrics, implement comprehensive testing strategies that reflect real world complexity, and meaningfully incorporate human feedback. You'll leave with a practical framework to confidently assess your agent's capabilities and ensure reliable performance when stakes are high.