Cover Image for Open Haus Berlin: An evening for developers building in the open
Cover Image for Open Haus Berlin: An evening for developers building in the open
Avatar for Together AI Calendar
Hosted By

Open Haus Berlin: An evening for developers building in the open

Register to See Address
Berlin, Germany
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

​An evening for developers building in the open.

What does it actually take to run open models in production for millions of users?

Engineers from n8n, NVIDIA, and Together AI share what they’ve learned building and operating real systems with open models. We’ll get into the practical engineering decisions behind model selection, evals, inference and serving, latency and throughput, reliability, and scaling workloads as usage grows.

Join us in a digital art museum for an evening of meeting like-minded builders, learning from engineers making open models work at scale in real production systems, good food and even better vibes.

Agenda
5:30 p.m. – 6:00 p.m. Arrival, drinks & networking
6:00 p.m. – 6:45 p.m. Three talks on building with open frontier models

  • ​Scaling context in production LLM systems: Riqwan Thahamir, Staff Engineer, n8n

    • ​As LLM applications grow more complex, managing context efficiently becomes critical to performance and cost. In this session, Riqwan Thamir, Staff AI Engineer at N8N, shares practical strategies from N8N's own infrastructure, including context engineering for multipurpose yet efficient agents, managing prefix cache, dynamic tool loading, and progressive context and skill loading. Attendees will walk away with concrete techniques for building systems that scale without sacrificing speed or reliability.

  • ​Agent-Aware KV Caching: Harry Kim, Principal Product Manager, NVIDIA

    • ​AI agents spawn sub-agents, compact their context and run long multi-turn sessions, so their KV cache patterns are hard to predict. This talk shows how the Dynamo router and the KVCR library work together to share KV cache across workers, recover it when an engine fails instead of recomputing it, and use signals from the agent harness to keep the KV that matters and evict the rest.

  • ​LLM inference for the era of agents

    • ​Zain Hasan, Staff AI/ML Engineer DX, Together AI

​6:45 p.m. – 8:15 p.m. Bites + conversations with fellow builders
8:15 p.m. Last call
8:30 p.m. See you next time

Location
Please register to see the exact location of this event.
Berlin, Germany
Avatar for Together AI Calendar
Hosted By