

How to Get AI to Score 90% in JEE Math Using RL (Aryabhata 1.0)
Aryabhata 1.0 is a 7B mathematics LLM that achieves 90.2% on JEE Main April 2025, outperforming OpenAI’s O4 Mini and Gemini Flash 2.5. In addition to model merging and supervised fine-tuning, the model leverages novel Reinforcement Learning with Verifiable Rewards (RLVR) using the A2C objective with group-relative advantage estimation. This talk will dive deep into the innovative RL training infrastructure, covering Adaptive Group Resizing and Temperature Scaling exploration strategies, as well as real-world reward hacking experiences and the mitigation techniques used to address them.
About the Speaker:
Sachin Dharashivkar is the CEO of AthenaAgent, an LLM post-training research lab, and a LossFunk Batch 2 resident. He collaborated with PhysicsWallah to train Aryabhata 1.0. Ritvik Rastogi is an ML Engineer at PhysicsWallah and a Co-Creator of Aryabhata 1.0.
Speaker’s socials • sachin-dharashivkar • sachdh • RitvikRastogi19