Avatar for Conscious Engines
Presented by
Conscious Engines
6 Going

Bangalore Paper Club - Edition #3

Register to See Address
Bengaluru, India
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

Researchers and builders rarely get together. That's exactly what The Research Room is for.

Every few weeks, we bring together a small group of researchers, engineers, and builders to work through influential ML papers; not just discuss them, but unpack the ideas, assumptions, and math that makes it amazing research.

catch up on Edition 2 here

Expect someone walking through a proof, reasoning through an architecture and having the kind of debates that only happen when everyone in the room has actually read the paper.

This edition is all about making models smaller without making them worse.

1. BitNet: Official inference framework for 1-bit LLMs (Microsoft) https://github.com/microsoft/bitnet

What if the weights were just -1, 0, and +1? BitNet is Microsoft's ternary line end to end: the b1.58 formulation, a 2.4B model trained on 4T tokens, and the kernels that make it run. Up to 6.17x faster on CPU with 55% to 82% less energy, and a 100B model running on a single laptop-class CPU.

2. CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs (Wang et al., ICML 2026 Oral)https://arxiv.org/abs/2606.26650

Ternary models have always required training that way from scratch. CAT-Q makes it post-training instead, ternarizing an existing checkpoint from just 512 calibration samples while beating BitNet b1.58 with roughly 100,000x fewer training tokens. It scales to 235B parameters in under 60 GPU-hours.

3. TurboQuant: Redefining AI efficiency with extreme compression (Google Research, ICLR 2026)https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/

Weights are only half the memory problem; the KV cache is the other half. TurboQuant pairs PolarQuant, which eliminates the per-block quantization constants everyone else pays for, with a 1-bit Johnson-Lindenstrauss correction, giving a 3-bit KV cache with no training and no accuracy loss. Attention logits get up to 8x faster on H100.
Format

Each paper is presented by an attendee (15–20 minutes), followed by an open discussion.

We will go beyond the abstract and discuss derivations, architectural decisions, experimental methodology, and what actually holds up in practice.

Small group, capped attendance, so everyone gets to challenge assumptions, ask questions, and explore ideas without the pressure of a conference Q&A.

Who should come

Researchers, ML engineers, graduate students, and builders who enjoy digging into papers beyond the headline results. Reading the papers beforehand isn't required, but you'll get much more out of the discussion if you do.

What to bring

A laptop or notebook, the papers (shared ahead of time), curiosity, and a willingness to change your mind.

Hosted by Conscious Engines, Bangalore.

Location
Please register to see the exact location of this event.
Bengaluru, India
Avatar for Conscious Engines
Presented by
Conscious Engines
6 Going