MMBU Challenge
About the Challenge
The MMBU Challenge, run out of the Stanford Medical AI and Computer Vision Lab, evaluates models through open-ended VQA on the Massive Multimodal Biomedical Understanding Benchmark (MMBU), an expert-curated benchmark of biomedical images.
Evaluation
Submissions are scored both on their answers and on the biomedical context surrounding each image — including modality, specimen, anatomy, and preparation.
Tracks
Three tracks cover frontier models, medically adapted models, and efficient models under 4B active parameters, allowing both large labs and small teams to compete on ground suited to them.
Resources & Prizes
Teams receive a public development set, weekly office hours with the organizers, special credits from Anthropic and GXL, and $20k+ in cash prizes. The top entry in each track will also be invited to contribute to the MMBU Challenge technical report.
The MMBU Challenge is supported by Stanford AI Lab, Anthropic, GXL, AWS, Highlanders, and Biohub.
