

MLn Club (ML Reading Group) #13: Manifold-Constrained Hyper-Connections
Welcome to Week 13: mHC: Manifold-Constrained Hyper-Connections
How does constraining residual connections to a geometric manifold restore identity mappings and enable stable scaling in large language models?
Can richer cross-layer connectivity improve reasoning and representation quality without sacrificing optimization stability?
The Paper Link Here
DeepSeek’s Manifold-Constrained Hyper-Connections (mHC) introduce a principled fix to recent attempts at widening residual streams. While Hyper-Connections improve performance by allowing richer cross-layer mixing, their unconstrained residual matrices break the identity-mapping property that makes deep networks trainable, leading to severe instability at scale. mHC resolves this by projecting residual mixing matrices onto the manifold of doubly stochastic matrices, ensuring residual updates remain convex combinations of features and preserving signal magnitude across depth.Empirically, this constraint dramatically stabilizes training in large language models, eliminating gradient explosions observed in unconstrained HC while retaining its performance gains. With careful systems co-design—kernel fusion, recomputation, and pipeline overlap—mHC incurs only ~6–7% overhead at 27B scale, while consistently improving downstream reasoning benchmarks. The result reframes architectural scaling as a geometric constraint problem, showing that richer connectivity can be safely exploited when paired with mathematically grounded structure rather than ad-hoc regularization.
Join us at CASI for discussion at 8 pm, and (optional) quiet reading from 7 pm.
📖 Reading Recommendations, Questions, or Comments? Contact us here!
🔎 View past meeting notes here.
What's this?
A super warm group of folks discussing their favorite topics!
In the first half, we host an optional quiet reading space
In the second half, we have a discussion where people can talk about what they found interesting about the reading and ask questions about things they didn't understand
When/Where:
CMU AI Safety Initiative's Office, 201 Craig Street, right across the PNC bank. Look for the open door up the stairs.
8pm discussion, 7pm optional quiet reading time.
Here's how it usually goes:
7:00 PM — arrival and settling in
8:00 PM — introductions
8:10 PM — discussion time
9:00 PM — wrap up then open discussion
Who's it for?
People who've been wanting to read up on the latest papers in ML and other fields but just haven't been able to find the time/motivation.
Why:
We've been procrastinating too much on our readings, even though we have so much fun doing them. We know we're not alone in this and want to keep others accountable for learning more about what they're passionate about!
We've also met a ton of really fun friends by discussing what we care about!
Rules/guidelines on how to act:
Act like a host, include people in conversations, talk to people even if they're strangers, offer to explain what you know, and keep an open mind! come to read stuff and find super fun friends :)
Bring snacks if you're feeling kind!