Cover Image for TKK #36 From Sound to Text: How Speech Recognition Really Works
Cover Image for TKK #36 From Sound to Text: How Speech Recognition Really Works
Avatar for wwktm
Presented by
wwktm
wwktm event calendar
Hosted By
34 Going

TKK #36 From Sound to Text: How Speech Recognition Really Works

Registration
Welcome! To join the event, please register below.
About Event

Speech recognition may seem like magic - you speak, and your words appear as text. In reality, the idea is quite simple. A model learns to convert speech into text by training on thousands of hours of human speech recordings, each matched with its correct transcript. By seeing these examples repeatedly, the model gradually learns to predict the correct words from new speech.

In this session, we'll try to unpack how ASR (Automatic Speech Recognition) models are actually trained and how they work when you use them, without diving into heavy math or academic theory.

We'll also go beyond audio-only models and look at large language models that also can listen. We will discuss why this matters, how the pieces fit together, and where this is headed next.

------------

Rupak Raj Ghimire (Google Scholar)
PhD Scholar, Information and Language Processing Research Lab (ILPRL), Kathmandu University

Rupak Raj Ghimire is a PhD scholar at ILPRL, Kathmandu University, researching Nepali Automatic Speech Recognition. He has over 15 years of experience in software engineering, working with C#, Python, and PostgreSQL. His research includes work on ASR fine-tuning, corpus development, and speech recognition systems for low-resource languages such as Nepali.

Location
Arbyte Solutions Pvt. Ltd.
Kathmandu 44600, Nepal
5th floor of everest bank / trikaya building
Avatar for wwktm
Presented by
wwktm
wwktm event calendar
Hosted By
34 Going