

FAQ Assistant in Practice: End-To-End Flow
Managing a growing technical community means answering hundreds of repeating questions across different channels every single day. In this session, we'll unpack the practical architecture behind the DataTalks.Club AI FAQ assistant, a production pipeline built to automate community support and deliver accurate answers directly where students ask them.
We will walk through the entire system lifecycle, moving from raw community knowledge sources to a real-time, deployed Slack assistant. You’ll see the exact data engineering steps used to gather unstructured community inputs and turn them into a searchable vector index, followed by a live demonstration of the bot running end-to-end inside Slack.
We'll Cover
Student question ingestion and FAQ dataset contributions
Multi-source data extraction across Slack threads and YouTube transcripts
Vector database indexing and search retrieval workflows
RAG system architecture and LLM response generation
Slack API integration and automated bot deployment
End-to-end live demonstration inside the active Slack community
By the end of this session, you’ll understand how to build and deploy a production-ready, multi-source RAG pipeline that turns unstructured Slack threads and YouTube transcripts into a responsive, automated Slack assistant.
We've also documented the full technical setup and step-by-step implementation in this article.
About the Speaker
Alexey Grigorev is the Founder of DataTalks.Club and creator of the Zoomcamp series.
Alexey is a software and ML engineer with over 10 years in engineering and 6+ years in machine learning. He has deployed large-scale ML systems at companies like OLX Group and Simplaex, authored several technical books, including Machine Learning Bookcamp, and is a Kaggle Master with a 1st place finish in the NIPS'17 Criteo Challenge.
DataTalks.Club is the place to talk about data. Join our Slack community!