

Bangalore Paper Club - Edition #2
Researchers and builders rarely get together. That's exactly what The Research Room is for.
Every few weeks, we bring together a small group of researchers, engineers, and builders to work through influential ML papers; not just discuss them, but unpack the ideas, assumptions, and math that makes it amazing research.
catch up on Edition 1 here - https://www.youtube.com/watch?v=1TROrlh_46E
Expect someone walking through a proof, reasoning through an architecture and having the kind of debates that only happen when everyone in the room has actually read the paper.
This edition, we're covering:
1. LLaDA: Large Language Diffusion Models
arxiv.org/abs/2502.09992
Can language models move beyond generating one token at a time? LLaDA introduces a scalable diffusion-based alternative that revisits how text is generated, opening up new possibilities for parallel decoding and efficient inference.
2. Nemotron-TwoTower: An Efficient Architecture for Diffusion Language Models
arxiv.org/abs/2606.26493
As diffusion language models become practical, what architectures best support them? NVIDIA's TwoTower design separates contextual understanding from denoising, demonstrating substantial inference speedups while maintaining strong language quality.
3. A Graph Talks, But Who's Listening? Rethinking Evaluations for Graph-Language Models (Petkar et al., Findings of ACL 2026)
https://aclanthology.org/2026.findings-acl.1624/
Today's Graph-Language Models look multimodal, but this paper shows they're quietly cheating, acing benchmarks using text or structure alone without ever fusing the two. The authors expose this with sharp behavioural and mechanistic analysis, revealing that plain soft-prompted LLMs keep pace with far heavier GLM machinery. CLEGR is a genuinely multimodal reasoning benchmark plus a data-scaling recipe that finally forces graphs and language to talk and listen.
4. Canine Olfaction Combined With Bayesian Modeling for Multicancer Detection From Breath Samples: A Phase II Study in India
https://ascopubs.org/doi/10.1200/JCO-25-02310
Trained dogs sniffed out seven types of cancer from just a breath sample across 3,275 people in Karnataka, India; no needles, no scanners. Paired with Bayesian fusion modeling, this canine olfaction system offers a low-cost, high-sensitivity triage test built for low and middle income settings. A wagging tail may be the future of accessible, non-invasive cancer screening.
Format
Each paper is presented by an attendee (15–20 minutes), followed by an open discussion.
We will go beyond the abstract and discuss derivations, architectural decisions, experimental methodology, and what actually holds up in practice.
Small group, capped attendance, so everyone gets to challenge assumptions, ask questions, and explore ideas without the pressure of a conference Q&A.
Who should come
Researchers, ML engineers, graduate students, and builders who enjoy digging into papers beyond the headline results. Reading the papers beforehand isn't required, but you'll get much more out of the discussion if you do.
What to bring
A laptop or notebook, the papers (shared ahead of time), curiosity, and a willingness to change your mind.
Hosted by Conscious Engines, Bangalore.