

Building AI Systems That Write Their Own GPU Kernels
What if AI could read research papers, implement the ideas, and optimize GPU kernels all autonomously?
This talk explores KernelEvolve, one of the four winners at the GPU Mode Hackathon 2025, an LLM-guided evolutionary system that can speed up production transformers.
We will dive into why evolution plus LLMs works when pure LLM generation fails, and ask whether AI systems can optimize their own infrastructure.
About the speaker:
Manoj Rao is a Fellow of Engineering at AMD’s AI Group, working on GPU optimization across current and future architectures. Previously, he worked on Tesla’s FSD and Optimus robots, led inference at Perplexity AI, and authored TorchServe, PyTorch’s model serving library. The project he will be discussing, KernelEvolve, was one of four winners at GPU Mode Hackathon 2025.
manojrajarao
http://github.com/manojrajarao
To attend online:
Add to calendar: https://bit.ly/3MAdNhL
Gmeet link: https://meet.google.com/gzk-ayys-smq?hs=122&authuser=0
Pre-read:
PyTorch Helion Blog, Google DeepMind AlphaEvolve Paper