

How to Build Smarter Optimizers: Beyond Just Minimizing Loss
Modern training often relies on extra losses to enforce desirable properties (e.g., L1 for sparsity, replay-based loss for preventing catastrophic forgetting).
As models grow in scale and the desired properties become more complex, directly enforcing such constraints becomes costly and inefficient.
Examples include preserving knowledge in class-incremental learning or reducing the tendency of the model to form unnecessary new feature spaces
In this talk, Subhash will cover learned optimizers and present his work on integrating property-based loss to create optimizers with built-in biases toward desired behaviors.
These optimizers are meta-trained on small models to internalize property-aligned update dynamics and then deployed at scale, where auxiliary losses become impractical.
The talk also hints at how the same idea could be used to discover stronger algorithms in areas like reinforcement learning.
About the speaker:
Padala SSSS V Sri Vishnu Subhash
B.Tech CSE’23 IIT Palakkad, C++ SDE at Arista Networks
subhashpadala
Preread:
[2507.12224] Optimizers Qualitatively Alter Solutions And We Should Leverage This
[1606.04474] Learning to learn by gradient descent by gradient descent
[2501.12670] Learning Versatile Optimizers on a Compute Diet
To attend online, please use the link below:
https://meet.google.com/drv-fwie-moi?hs=122&authuser=0
To add the event to your calendar:
http://bit.ly/3LZiH7M