Cover Image for vLLM Meetup Toronto
Cover Image for vLLM Meetup Toronto
Avatar for vLLM Meetups and Events
Join the vLLM community to discuss optimizing LLM inference!

vLLM Meetup Toronto

Register to See Address
Toronto, Canada
Registration
Registration Closed
This event is not currently taking registrations. You may contact the host or subscribe to receive updates.
About Event

​Join us for the vLLM x Cohere meetup in Toronto, an evening for AI engineers, researchers, and infrastructure builders.

​This is a chance to connect with vLLM maintainers, the Cohere team building and serving models for production-scale workloads, and engineers working on inference at NVIDIA. We'll cover how the vLLM ecosystem and Cohere's open-weights work and upstream contributions fit together.

​The session goes deep on inference, including the vLLM roadmap, speculative decoding, serving for agentic workloads, and scaling agentic RL. We'll close with a live panel on where inference is heading and what's next for open source AI, then finish the evening with networking, drinks, and bites.

​Speakers

  • ​Roger Wang, Co-founder, Inferact; Core Maintainer, vLLM

  • ​Ekagra Ranjan, Member of Technical Staff, Cohere

  • ​Sungjin Hong, Machine Learning Engineer, Cohere

  • ​Zhanda Zhu, System Software Engineer, NVIDIA; CS PhD, University of Toronto

  • ​Kosseila Hd, Lead Forward Deployed Engineer, Tensormesh

​Schedule

  • ​6:00pm - Doors open

  • ​6:30pm to 8:00pm - Talks and panel Q&A

  • ​8:00pm to 9:00pm - Networking reception

Location
Please register to see the exact location of this event.
Toronto, Canada
Avatar for vLLM Meetups and Events
Join the vLLM community to discuss optimizing LLM inference!