Democratizing training at scale: How to Train Stable Diffusion for $2,000!
In this talk, the BuzzRobot guest — Vikash Sehwag, a research scientist from Sony AI — will discuss his recent work on democratizing the development of large-scale text-to-image diffusion models.
As scaling laws push the capabilities of these models, they also simultaneously increase the computational cost, concentrating the training of these models among high-resource teams.
As a solution to this bottleneck, Vikash will present his approach to training high-performance large-scale diffusion models with an order of magnitude lower computational cost than existing state-of-the-art models.
He will introduce a patch masking technique, which, combined with a lightweight patch-mixer, alleviates performance degradation with masking.
Our guest will also discuss the effect of dataset choices, effect of input masking in diffusion transformers and CNNs, and the impact of other architectural design choices on enabling micro-budget training of diffusion models at scale.
Read the paper
Join BuzzRobot Slack to connect with the community