Cover Image for Paper Club: Memory, Modularity, and Safety: Memento as a Case Study in Safe and Performant AI Systems
Cover Image for Paper Club: Memory, Modularity, and Safety: Memento as a Case Study in Safe and Performant AI Systems
10 Went

Paper Club: Memory, Modularity, and Safety: Memento as a Case Study in Safe and Performant AI Systems

Registration
Past Event
Welcome! To join the event, please register below.
About Event

Technical Note: This event is intended for participants with a technical background. We strongly encourage reading the paper ahead of time to fully engage with the discussion.

Recent AI progress has leaned on ever‑larger monolithic LLMs, where core intelligence functions like memory, reasoning and planning are densely entangled, resulting in reduced interpretability and control.

In this week's paper club, we explore an alternative trajectory: constructing intelligence through modular architectures that separate concerns (planning, execution, memory, retrieval policy) to expose natural governance and safety levers, without compromising on performance.

Our focal case study is "Memento: Fine-tuning LLM Agents without Fine-tuning LLMs" (https://arxiv.org/abs/2508.16153), a planner-executor agent with an explicit episodic Case Bank and a tiny (2‑layer MLP) learned retrieval module that optimizes which past cases to surface rather than fine‑tuning the underlying LLMs. In this system, we can see episodic memories (past experiences) in plaintext, and trace its process of storage and retrieval which influences its decision-making process.

Despite its minimal parametric adaptation, Memento is ranked #1 in the GAIA validation set and #4 on the GAIA private test set, beating other agents like Manus. Memento also sets a new SOTA on SimpleQA, and ranks #2 on Humanity's Last Exam.

While Memento itself was not conceived as a safety project, it demonstrates how competitive performance can coexist with architectural choices that enhance transparency and control. We will discuss the architecture and modules at a high level, as well as potential threats to validity. If time permits, we can also discuss on other modular architectures, how it ties in with Memento, and opportunities in this line of research.

Attendees are highly encouraged to read the paper at a high level (can skip the math)

The paper can be found here: https://arxiv.org/abs/2508.16153

Location
Lorong AI (WeWork@22 Cross St.)
10 Went