Cover Image for Visual Intelligence Hackathon (Real Time Video Anomaly Detection)
Cover Image for Visual Intelligence Hackathon (Real Time Video Anomaly Detection)
Avatar for AI Hackers Collective
Where AI practitioners level up from AI-assisted to AI-native operators.

Visual Intelligence Hackathon (Real Time Video Anomaly Detection)

Registration
Welcome! Please choose your desired ticket type:
About Event

A car stopped in a parking lot? Normal.

The same car stopped on a highway? Now you probably want to know.

That tiny difference is exactly where most computer vision systems start struggling.

On 5th September, we’re bringing together AI/ML engineers, researchers, students, and multimodal AI builders for a full day to work on one problem:

Can we teach a small vision language model to spot what actually matters in live drone video, in real time?

And no, this isn’t a “prompt something cool and demo it at 6 PM” kind of hackathon.

You’ll be working with real drone footage, limited GPU capability, and the kind of messy visual context that makes this problem genuinely hard.

The Challenge

A drone flying over a city might see thousands of perfectly ordinary moments.

And then there’s the one that matters.

A vehicle has broken down on a highway. Or traffic is slowly building up. Or Smoke appears where there shouldn’t be any or anything unusual.

The difficult part is that an object itself usually isn’t the anomaly. The context is.

Traditional object detectors can tell you what they see. The challenge is building something that can understand whether what it sees is unusual enough to require attention.

Vision-language models can reason about that context, but large models are too slow and expensive to continuously run across live video feeds.

So we’re making the problem harder:

Make it work with a small model. Make it work in real time. And make it economical enough that it could eventually run across many drone feeds at once.

What You’ll Spend the Day Doing

We’re not dropping a problem statement at 9 AM and leaving you alone with Stack Overflow.

The morning starts with a state-of-the-art session on video anomaly detection, related work across robotics and computer vision, and demos from the FlytBase team.

Then we build.

You might fine-tune a small vision-language model, distill a larger model, build a lightweight detection + verification pipeline, implement recent research, or try something none of us thought of.

The approach is open.

The constraint is what makes it interesting.

What You’ll Get

  • A state-of-the-art deep dive before you start building, so you know what is actually possible today.

  • Real urban drone footage, including challenging conditions such as night flights.

  • Public benchmark datasets prepared in advance, so you spend the day building instead of downloading datasets.

  • A genuinely difficult ML problem involving video understanding, contextual reasoning, efficient inference, and open-world anomaly detection.

  • A room full of people working on the same hard problem, which means plenty of comparing approaches, breaking things, borrowing ideas, and learning from each other.

  • A full day to actually build and train, rather than sitting through talks about what AI might be able to do someday.

The complete problem statement, evaluation criteria, and submission format will be revealed on the day.

Who Should Join?

This is for you if you’re working with, learning, or seriously curious about:

Computer Vision · Multimodal AI · Vision-Language Models · Video Understanding · Model Fine-Tuning · Distillation · Efficient Inference · Anomaly Detection · Open-Set Recognition

AI/ML engineers, researchers, students, and builders are all welcome.

You don’t need to walk in knowing the answer.

You should walk in ready to code, experiment, train, fail a few times, and keep going.

One important thing:
arrive with your coding setup and model access already tested. If you plan to fine-tune, have your training environment ready too.

We’d rather spend the day fighting the actual problem than fighting CUDA.

The Day

9:00 AM – 9:30 AM
Breakfast + meet the people you’ll be building alongside

9:30 AM – 11:00 AM
State of the Art: Video Anomaly Detection + related robotics/CV work + FlytBase demos

11:00 AM – 6:00 PM
Build, train, test, break things, try again

6:00 PM – 7:00 PM
Selected demos + results

Date: 5th September 2026
Time: 9:00 AM – 7:00 PM

If you spend your weekends reading papers, fine-tuning models, experimenting with vision systems, or wondering what happens when multimodal AI leaves the benchmark and meets messy real-world video...

you’ll probably want to be in this room.

Register and come build with the AI Hackers Collective.

Join the whatsapp group for all the updates: Join WhatsApp Group

Location
FlytBase Labs
701, opp. Croma - Baner, Baner, Pune, Maharashtra 411045, India
Avatar for AI Hackers Collective
Where AI practitioners level up from AI-assisted to AI-native operators.