Cover Image for Backchannel Berlin (with Inworld)
Cover Image for Backchannel Berlin (with Inworld)
Avatar for Inworld Events
Presented by
Inworld Events
Inworld AI is a research lab building realtime speech models and inference for consumer applications.
Private Event

Backchannel Berlin (with Inworld)

Register to See Address
Berlin, Germany
Registration
Approval Required
Your registration is subject to host approval.
Welcome! To join the event, please register below.
About Event

​A research discussion on building frontier AI models, for people working on speech, audio and language models.

​End-to-end speech-to-speech models can now listen while they talk, handle interruptions and keep a consistent voice. Getting them to frontier quality is still largely open research: data, architecture, training objectives and evaluation. Running them for consumer apps, in every language and within a few hundred milliseconds, adds constraints of its own. The afternoon covers both.

​Open questions: speech-to-speech and conversation

  • ​What should a full-duplex model predict, and at what granularity, to handle turn-taking, interruptions and backchannels?

  • ​Cascaded STT, LLM and TTS versus end-to-end speech-to-speech: where does each break down outside benchmarks, and what closes the gap?

  • ​How do you keep an LLM's reasoning and instruction-following when you train it to listen and speak?

  • ​What data and speech representations does full-duplex training need, and how far can synthetic dialogue go?

  • ​How do you train a model to follow plain-language direction on delivery while keeping a speaker's identity stable?

​Open questions: at consumer scale

  • ​How far can cross-lingual transfer take these models in languages with little transcribed speech?

  • ​Which of batching, caching, quantization and speculative decoding matter most when a conversation has to respond within a few hundred milliseconds?

  • ​MOS, WER, speaker similarity and preference arenas each miss something. What should we measure for conversations, and how do we evaluate languages with few native raters?

​Format

​Our research team opens with 30 minutes on how we're building speech-to-speech and serving it in realtime, including what hasn't worked. After that, anyone can take 5 to 10 minutes to present their own work or a problem they're stuck on (tell us when you register), followed by open discussion. If you'd rather listen, you don't need to prepare anything. With speakers' agreement, we'll share slides afterwards. Food and drinks throughout.

  • ​15:00 - Doors and food

  • ​15:30 - Inworld research: building and serving speech-to-speech

  • ​16:00 - Short talks from attendees, and discussion

  • ​17:30 - Food and open discussion

  • ​19:00 - Drinks

​Who it's for

​Researchers and engineers working on speech-to-speech and full-duplex models, TTS, ASR, audio codecs, conversational modelling, LLM serving or evaluation, in academia or industry. PhD students are welcome.

​About Inworld AI

​Inworld AI is a research lab building realtime speech models and inference for consumer applications. We're working on full-duplex speech-to-speech and recently acquired Ultravox to accelerate that research. Our models team also trains our TTS and speech recognition models. TTS-2, our current TTS model, takes direction on delivery in plain language. Our inference research team works on backends, kernels, quantization and decoding. In production we serve trillions of LLM tokens and tens of billions of characters of speech a month.

Location
Please register to see the exact location of this event.
Berlin, Germany
Avatar for Inworld Events
Presented by
Inworld Events
Inworld AI is a research lab building realtime speech models and inference for consumer applications.