

Backchannel Berlin (with Inworld)
A research discussion on building frontier AI models, for people working on speech, audio and language models.
End-to-end speech-to-speech models can now listen while they talk, handle interruptions and keep a consistent voice. Getting them to frontier quality is still largely open research: data, architecture, training objectives and evaluation. Running them for consumer apps, in every language and within a few hundred milliseconds, adds constraints of its own. The afternoon covers both.
Open questions: speech-to-speech and conversation
What should a full-duplex model predict, and at what granularity, to handle turn-taking, interruptions and backchannels?
Cascaded STT, LLM and TTS versus end-to-end speech-to-speech: where does each break down outside benchmarks, and what closes the gap?
How do you keep an LLM's reasoning and instruction-following when you train it to listen and speak?
What data and speech representations does full-duplex training need, and how far can synthetic dialogue go?
How do you train a model to follow plain-language direction on delivery while keeping a speaker's identity stable?
Open questions: at consumer scale
How far can cross-lingual transfer take these models in languages with little transcribed speech?
Which of batching, caching, quantization and speculative decoding matter most when a conversation has to respond within a few hundred milliseconds?
MOS, WER, speaker similarity and preference arenas each miss something. What should we measure for conversations, and how do we evaluate languages with few native raters?
Format
Our research team opens with 30 minutes on how we're building speech-to-speech and serving it in realtime, including what hasn't worked. After that, anyone can take 5 to 10 minutes to present their own work or a problem they're stuck on (tell us when you register), followed by open discussion. If you'd rather listen, you don't need to prepare anything. With speakers' agreement, we'll share slides afterwards. Food and drinks throughout.
15:00 - Doors and food
15:30 - Inworld research: building and serving speech-to-speech
16:00 - Short talks from attendees, and discussion
17:30 - Food and open discussion
19:00 - Drinks
Who it's for
Researchers and engineers working on speech-to-speech and full-duplex models, TTS, ASR, audio codecs, conversational modelling, LLM serving or evaluation, in academia or industry. PhD students are welcome.
About Inworld AI
Inworld AI is a research lab building realtime speech models and inference for consumer applications. We're working on full-duplex speech-to-speech and recently acquired Ultravox to accelerate that research. Our models team also trains our TTS and speech recognition models. TTS-2, our current TTS model, takes direction on delivery in plain language. Our inference research team works on backends, kernels, quantization and decoding. In production we serve trillions of LLM tokens and tens of billions of characters of speech a month.