Cover Image for Vision AI for Robotics using ROS 2 (Cohort 3)
Cover Image for Vision AI for Robotics using ROS 2 (Cohort 3)
Avatar for Packt Publishing
Presented by
Packt Publishing

Vision AI for Robotics using ROS 2 (Cohort 3)

Virtual
Get Tickets
Ticket Price
$119.99
Welcome! To join the event, please get your ticket below.
About Event

Overview

Build AI-Powered Robot Perception using Vision Language Models (VLMs)

Vision Language Models (VLMs) are revolutionizing robot perception by enabling robots to understand and reason about their surroundings using both visual and natural language inputs. In this hands-on workshop, participants will learn how to integrate lightweight VLMs with ROS 2 to develop intelligent robot perception applications that run efficiently on standard laptops, with or without an NVIDIA GPU. The workshop will also explore GPU acceleration using NVIDIA CUDA to optimize inference performance for real-time robotic applications.

Meet Your Instructor

Lentin Joseph — Author of 11 ROS & Robotics books, Co-Founder of RUNTIME Robotics, and TEDx Speaker.

By the End You'll Walk Away With

  • A complete ROS 2 AI Vision application in a robot

  • Hands-on experience with Vision Language Models (VLMs)

  • ROS 2 image processing pipeline using OpenCV

  • Image Captioning and Visual Question Answering (VQA)

  • ROS 2 Services for Vision AI inference

  • Knowledge of CPU deployment and NVIDIA GPU acceleration

Who Should Attend

  • ROS 2 Developers

  • Robotics Engineers

  • AI & Computer Vision Engineers

  • Students and Researchers

  • Anyone interested in intelligent robot perception

What You'll Learn

  • Introduction to Vision Language Models

  • ROS 2 camera and image pipeline

  • Running lightweight VLMs on CPU

  • Image Captioning

  • Visual Question Answering

  • ROS 2 Topics, Services and Parameters

  • Performance optimization

  • Optional NVIDIA GPU acceleration

What You'll Need

  • Laptop (Ubuntu 24.04 recommended)

  • Python 3.10+

  • ROS 2 Jazzy

  • USB Webcam

  • Minimum 8 GB RAM (16 GB recommended)

  • Internet connection

Why Now?

Vision Language Models are becoming a key technology for next-generation robotics and Physical AI. Learning how to integrate them with ROS 2 gives developers practical skills for building intelligent robot perception systems used in research and industry.

🎟️Reserve your seat now - Limited seats. Live support. Real builds.

*By signing up for this event, you agree to receive emails from Packt Publishing.

Avatar for Packt Publishing
Presented by
Packt Publishing