

Hackathon Winners Present: Do Multilingual Vision-Language Models Abstain under Cross-Modal Conflict in Low-Resource Languages?
A model sees a red car. The caption says blue. Which does it trust?
In English, it sometimes catches the conflict. In Hindi and Telugu, that safety check breaks down.
Anvesh Reddy Lankala, Vicky Feliren, and Akansh Jain built a benchmark across English, Hindi, and Telugu and tested 9 vision-language models. Qwen2.5-VL-7B looked safe on standard metrics — but overrode correct visual evidence 81% of the time when text disagreed.
Winners, Asia-Pacific track, Apart Research's Global South AI Safety Hackathon.
Hosted by AI Safety India, Weekly Discussion Series.
Format:
- Research presentation
- Q&A, open floor
- AI Safety Mixer
Speakers
Anvesh Reddy Lankala
Vicky Feliren
Akansh Jain
Full submission: https://apartresearch.com/project/do-multilingual-visionlanguage-models-abstain-under-crossmodal-conflict-in-lowresource-languages-hcsp?utm_source=luma