Cover Image for AI Tuesdays: Research Paper Reading
Cover Image for AI Tuesdays: Research Paper Reading
20 Went

AI Tuesdays: Research Paper Reading

Hosted by Amrut Rajkarne, Varad Maniyar & Nexus VP
Register to See Address
Bengaluru, India
Registration
Registration Closed
This event is not currently taking registrations. You may contact the host or subscribe to receive updates.
About Event

Most of what we know about where AI is going arrives as a paper first. But papers are read alone and rarely argued about with anyone who has read them closely.

What changes when a room of researchers and builders works through the same paper together?

For our next edition of AI Tuesdays, we're running a research paper reading. A small group, a few papers, each presented by someone in the room, followed by open discussion. The list of papers in this reading session are attached below.

What we'll get into

  • The method, not the abstract. How the result was actually produced, and what the setup takes for granted

  • What holds up. Which findings survive contact with a different dataset, a different scale, or a production environment

  • What the baselines hide. Comparison choices and the gap between reported and practical performance

  • Why this paper now. What it changes for anyone building on top of it over the next year

Format

  • Each paper is presented by an attendee for 15-20 minutes, followed by open discussion. List of Papers attached below.

  • Doors open at 7:00 PM; the first paper starts at 7:15 PM sharp.

  • Reading papers beforehand isn't required, but highly suggested, so that everyone is upto speed.

  • Food and beverages will be served

Paper #1:
Recursive–Structured-State–Termination (RST) - Improving the performance of RLM for long-context reasoning with the help of structured graph state

Transformer-based language models achieve strong performance across NLP tasks, but their effectiveness degrades in very long-context settings due to context rot. Although recursive language models mitigate this limitation by processing inputs in chunks, they introduce a control challenge in determining when to continue reasoning, consolidate evidence, or terminate. We propose the Recursive-Structured-State-Termination (RST) Reasoning Engine, an adaptive recursive framework that incrementally extracts knowledge, maintains a structured reasoning graph of intermediate states, and employs convergence-based termination to improve reasoning efficiency, transparency, and scalability.

Details on the paper at this link.

Paper #2

What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Engine

We ask what gets a language model onto the Apple Neural Engine (ANE) and what makes it fast there, and we answer with three measurements. We sweep a 64-shape matrix of LLM primitives that varies how a computation is expressed while holding what it computes fixed, recording per-operation device support. We then train matched models across size and precision, with quantized checkpoints byte-identical in structure to their fp16 counterparts, so every deployment measurement is of a real trained artifact. And we read the ANE's memory-controller byte counters during inference, establishing what actually ran rather than what the compiler intended. We support every headline claim with at least two of these three measurement paths. We find that placement is a property of how a computation is expressed, not of what it computes: a fused RMSNorm is fully ANE-eligible while its arithmetically identical decomposition is CPU-only.

Details on the paper at this link.

Paper #3

Escrow: WHEN A RECORD STREAM HAS EARNED A NEW NODE

Each arrival in a stream of records raises the same question: does the record join a node that already exists, or has the stream earned a new one. The rules in use answer with a constant chosen by hand: a distance cut-off in streaming clustering, the penalty the small-variance limit leaves in DP-means and BP-means, or the sampling temperature of a language-model pipeline. In the second case the constant is what a deletion leaves behind. Before the limit is taken the new-node penalty is a code length, the limit erases the part of it that could have been computed, and we do not take the limit. We read the price of a node off a code that predicts each record before seeing it, so the price is exact and grows with the logarithm of the stream, and no quantity in the rule is calibrated on data. No fresh node can pay that price on its first record under any valid code, so evidence accrues in escrow against a node that does not yet exist and the node is created when its account covers its price. On pure noise it creates nothing, where streaming decision trees grow to a mean of ninety-one false nodes, and it recovers planted structure exactly in all twenty arrival orders. On raw Wikipedia infoboxes its agreement with the category labels runs from 0.64 to 0.99 with the record order, which is the measured cost of a one-pass greedy sequence, and a language model given the same records is more accurate there but returns a different graph on every seed.

The Full Paper is attached.

New Node Paper.pdf
552.8 KB

There might be a couple more presentations about other Researchers building frontier models

Who it's for

AI researchers, ML engineers, graduate students, and anyone who reads papers closely and wants to argue about them with people who do the same.

This is an invite-only dinner event with 10-12 attendees.

Hosted by Amrut Rajkarne, AI Investor at Nexus Venture Partners.

Location
Please register to see the exact location of this event.
Bengaluru, India
20 Went