"Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"

Tetherless World · Intermediate ·🧠 Large Language Models ·9mo ago

Key Takeaways

The video discusses full-duplex modeling with audio codec LMs for fluid AI conversation with natural turn-taking, highlighting the importance of paralinguistics and incremental speech processing in achieving human-like dialogue.

Full Transcript

You want me to see? So, welcome everybody to our uh our second TWED uh of the fall 2025 season. Um let's hear for TW. Let's hear for Abraham Sander. going to be giving us a great talk. Uh just a quick couple logistics. We are recording this. It goes down to the web forever. Um uh you want allowing questions during your talk or would you like to get to it? >> I'd say I'll I'll pause at specific moments for questions. >> Okay. >> Because there is a lot of content and I want to make sure that we have a chance at getting through it. and um and our our next wed is going to be Enrique uh in three weeks uh on the 22nd. Um so without further ado, take it away to Abraham. Thank you. Also, >> welcome everybody to uh the TWED talk. Um it is titled Fluid AI conversation with natural turn taking full duplex modeling with audio codec LMS. And hopefully by the end of this talk you'll have some idea what that means. I'm still figuring out what it means because it's the topic of my dissertation. Um, okay. So, we'll talk about um we'll introduce the the topic. We'll talk about how turn taking works and real humanto human dialogue. Um then we'll go a little back background on spoken dialogue agents and then how we get into the techn technicalities of how we might want to go about building a full duplex dialogue agent that is humanlike the codec language model. So if you have two people who are talking to each other, you have a lot of rich um interaction potential that beyond just the content of what they're saying. laughter, you have pausing, um specific coordination of turn taking, um interruptions, back channels, acknowledgements like yeah, uh you have filled pauses such as uh or um these things, like if you ever talk to chat GPT, you notice that uh in their audio mode, all of that stuff is completely absent. It's just like talking to a a very like a studio quality robot. So um yeah all of these things um I call them paral linguistics well as other people call them that as well but the non-contentful content of what you're saying actually does hold a lot of content. It holds signaling and coordination information um that lets a speaker know whether or not you follow them or whether you're engaged with them or whether you're listening to them or um whether they whether you need to go back and clarify something etc. So there are there have been um lots of history of this stuff being studied both in cognitive science and computer science. Um in uh specifically cognitive science there is a field called communication um or or um sorry uh uh conversational analysis or CA um not not communications. Conversational analysis involves analyzing transcripts of spoken conversations and um annotating at what points in those conversations these paral linguistic phenomena occur and then being able to study frequency statistics and other things to to get a sense for how humans communicate. So traditionally spoken dialogue systems that like chat GT are implemented as cascaded models. So you have some speech recognition system. You might have an LLM in the middle that's acting as a chatbot. And then you have a texttospech system that is transcribing the responses into speech. But these models fail to produce all of those power linguistic phenomena because and and and thus they're not fluid and lielike. They it's like talking to a studio quality robot. So why is this? And one reason is because they they they uh implement rigid um half duplex turntaking. Half duplex specifically means that only one person can talk at a time. There isn't a parallel birectional communication which would be full duplex. And also there's no modeling of the protic and paral linguistic features that we just talked about. modeling of laughter um and of rate and pitch variation and filled pauses and things like that that you might expect in a normal conversation and also cascaded systems suffer from an inescapable issue which is error propagation. So if if there is a um an error that is made by the speech recognition system that will propagate to the LLM chatbot which has never actually heard the audio. It's just received some text is treated as gospel from the speech recognition system and that same thing the text to speech receives some text from the chatbot treats that as gospel has no idea whether or not it is actually the result of some kind of an error. So the further back in the pipeline that the error happens the more can go wrong later on. So like if you say what's going on and here's what's watch going on and oh is that a show and be like well that doesn't sound coherent at all given what I asked. So um so why not handle all three of these functions in a unified model and specifically a full duplex auto reggressive model and that's what we're going to explore today in this talk. And again full duplex means that each party in the conversation has the ability to speak simultaneously. there is a true birectional communication that is open at all times. So both parties have to figure out when is the appropriate time to speak and when is the appropriate time to yield otherwise they'll just talk over each other. So I'm going to go through this section a little bit fast. Um this is the turn taking in human dialogue is just just a little bit of a background on what's been studied in the cognitive science literature on this. So there was a a seminal paper by Sax, she and Jefferson in 1974 um and other follow-on papers that discuss basically create a systematics for analyzing human dialogue. Um and there are are essentially units that can be taken such as the turn constructional unit, interpausal unit and transition relevant transition relevant place which determine um how you can compose a a duplex human conversation. So for example, a turn constructional unit is going to be pos potentially multiple chunks of audio separated by a pause. Whereas an inter interpausal unit is what you can use to construct your turn. Multiple things that you say sequentially with maybe some pause in the middle. Then you have back channel acknowledgements which are small interruptions like okay or uhhuh which are not meant to take the floor from the other speaker but are there to help guide the conversation by telling the other speaker that you're following what they have to say um and again you have all of these different possible phenomena. You have positive and negative offsets. So either a g uh either a a gap, a pause or an overlap, meaning people over each other. Filled pauses such as um and you know, which are specifically there in conversation to signal to the other speaker that you're not done speaking. Don't start talking yet. I say I just have to think about back channels. Like I was just saying, they help tell the other speaker that you're following them or maybe not. Um and then you have other prosotic clues such as intensity of the speech, speed of the speech, emphasis on words, tone, drawn out syllables, etc. Um you have laughter, you have singing, grunting, squealing, humming, singing, all of these things that can occur in within human speech. And uh you also have things like coral speech, speaking in chorus, such as ever heard two people greeting each other at the same time, hi, how are you? while the other person's also saying, "Hi, how are you?" at the same time. So, cascaded half duplex spoken dialogue systems that kept don't do any of these things well at all. Um, so this is a transcript from an actual human dialogue shared on um on TalkBank, which is a repository of transcribed phone calls. Um, let's see if I can play one of these. Make me log in. Okay, give me a second. >> Okay. Set up the m music. >> What? >> Set up the music. >> Why don't you go first? Um, >> I guess I'd have to raise my my computer volume to around doesn't work. If I wanted to share audio to be played on >> Oh, I don't I I don't know how that work. >> You probably want to increase the volume of the audio from the TV, right? Is there a remote somewhere? I mean, it's coming from me. >> Doing that, but we don't hear anything. >> Oh, I know what to do. >> Well, >> I don't know. I just been going crazy lately. Like, sometimes it seems like every Sunday is like a real like dollar for me. Not like it's like a big like loser day for me. Like I had this big paper to write. >> Oh, that's what I'm doing right now. >> And I left it to the last minute. Like I was in class when we had to choose our poet on like Wednesday. last Wednesday and my professor goes, "So, who are you choosing?" And I go, "Ezra Pound." Like right off the bat, you know what I mean? It wasn't even on the list. I just wanted to like like Ezra Pound is like pretty big. It's like a big thing. You know what I mean? So, my professor was all like winking and smiling at me. He was so pleased. So, like yesterday, I go to the bookstore and I buy Ezra Pound poetry. And I come home and I'm like sitting in the bathtub trying to read this and I go, "What?" Like I've never encountered poetry like that before. It was like her span down the last and then I cumin Greek statement. >> So I think you that's that dialogue up there is you. Okay. Oh, that doesn't do anything if I click on it. Um >> I don't know. I think this got a >> okay >> some sort of >> I've unmuted myself >> or something. It wasn't like poetry for the masses at all. >> So that day borders >> and like you know I couldn't buy like three of Sonia Sanchez's books. You know what I mean? That's like 40 bucks, >> right? So I just sat there for like three hours and copied down 30 poems and then I like walk home and I'm like all depressed and I'm like writing. >> So this is coming out of my speakers. >> All this stuff like nose grows all this just like all this sad depressing stuff in my life and I think I should write a poem. I should write a poem. And I just I was walk >> I did. at home like all depressed and sad and blue and then I just started typing all the you have this like orange color pad in the top of your like information it won't show Yes. >> Okay. There should be three buttons for example. Go there and there should be share computer. >> That's already checked. >> I don't know. >> Just keep going. >> All right. I'm sorry, guys. Um Um >> You can't play audio just from your laptop, can you? I I can, but you guys aren't going to hear it unless you over here. >> Let's keep going. >> All right. Well, we tried. Um, so I'll I'll go back to the slides. So humans use turn taking cues beyond timing of interpausal units to predict when or when not to take their turns. Um they use other those other cues include um the content of the utterance such as syntactic completeness whether or not you actually finished what you had to say. Uh pro proity uh whether you're holding a pitch or whether it's increasing or decreasing. Um and the average uh gap length that that occurs in natural human dialogue has been measured in specific corpora um to be about 200 milliseconds. So about 200 milliseconds on average will go by between when one person starts stop speaking and the next person starts. And this indicates that people have the ability to predict when somebody's turn will end and and when they would be able to come in and start responding. and that they can prepare part of that respond response in advance. Um, spoken dialogue systems don't you really work this way. They don't use those type of of acoustic cues and they don't predict turn taking events in advance. rather they uh typically use a voice activity detector to predict when this the speech has stopped and then start uh queuing up a response at that point using whatever LLM component is is part of the system. Um if the dialogue system was able to create an overlap it might want to support these types of overlaps which are measured or which are observed in human. One is called competitive overlap, which is where maybe one person is starting to to speak when the before the other person is done and then maybe one of them will stop speaking or continue speaking to signal, hey, I'm not ready for you to speak yet. Um, and it might sound like an interruption. Um, then there's cooperative overlap such as um back channels, which we discussed, things like mhm, uh-huh, and uh conditional access overlaps, which is like a sentence completion. If you have somebody who is able to predict what you're going to say and then they come in and then they complete your sentence for you or how that happen um permanent overlap happens um oops sorry actually I'm not sure I forget exactly what terminal overlap is I probably I'm thinking of a different word for Um this is on a slide which I yeah I don't remember what it means. So moving on to computational models of turn taking um with the goal of of improving the fluidity of spoken dialogue systems. So work has been done in this area for many many years going all the way back to 200 or the the early 2000s. The um neural models of turn taking have been explored since about 2009. Um well, if there's any anybody who actually knows this field in the room, they're going to be saying, "What are you talking about?" Neural models have been explored a lot earlier than this, but the first um operationalized turntaking model that was were basically an attempt to create a a duplex dialogue system with it was in 2009. This was the number system which uh basically just accepted a person dictating a stream of numbers and then um it would respond to confirm whether or not it would accurately whether it accurately heard what the stream of numbers is and could interrupt you if it felt like it uh had not heard numbers correctly for a confirmation. Um you have uh sconcy in 2007 2017 presented a uh a predictive model of turn taking. So given a stream of audio it could predict at what points the uh a turn switch would happen and it used a a RSTM for that. Um similarly meer Hugh and Schlangan in 2017 built a neural model that used both acoustic and lexical features. So combination of text and uh and audio. Um and then there was similar work in 2019 where there was a um a probabilistic model of when a transition relevant place would happen. Transition relevant place is not specifically when a turn happens but when a turn could happen. So it's predicting all of the places where a turn switch would be reasonable in the dialogue. And here are some transcript based models that use just text to try to do the same thing. Essentially, if you can transcribe the spoken dialogue, then the uh then you can potentially guess where a turn might be taken or where place might even without the audio. So now, how would we operationalize this as a spoken dialogue agent? So this section just goes a little into what spoken dialogue agents are and uh how they're traditionally implemented. So if you ever had one of these, it's an Amazon Echo, right? Um it is like most other dialog systems. Well, I don't think people really use them anymore. or at least I've never seen them recently, but I think they still exist. Like most dialogue systems, an echo is um a cascaded dialogue system. It has a speech recognition component, um an LLM or some kind of natural language NLP processor in the middle and a texttospech system. Um it is non-incremental which means that um all of the steps of understanding what was said and speaking and choosing what to say next is happening in series in actual uh human conversation and in a duplex. We would want this to happen in parallel. So you could be listening to what's being said, understanding what's being said, getting ready to respond all while um the all at the same time essentially. So the number system from Scotty and Schlongan was the first system to do this successfully even though the domain of the conversation was just streams of um spoken numbers rather than general conversation. Um, since we're having trouble with audio, I'm going to skip this. We'll we'll come back and try to play some of the later audio if we can, but I'm going to skip this. Um so incremental speech processing which is that model where like in humans and in duplex dialogue systems um you have understanding and planning for for for response all happening at the same time. Turn taking is done continuously. You don't the other person's done speaking to decide whether you want to to take your turn. You are con constantly every couple of milliseconds predicting whether or not now might be a good time to butt in. So the question then is well why can't we handle all of these functions in one auto reggressive model that will retain both the acoustic and semantic context at all stages and can that model also be a continuous model of uh turn taking as well. So that brings us to codec language models which is the tools that we would use to implement such a system. So a back quick background on language modeling which I think none of you in this room need but it's here. Tokens are generated auto reggressively after being tokenized. So you have texted in text is broken into subword units called tokens and they are they are uh converted into um embedding indices which are then looked up in an embedding matrix. The embedding matricy matrix uh returns a vector for each token which represents its sort of semantic uh position within the space of possible tokens and that sequence of tokens is then fed into the transformer. Um and the transformer is able to use attention to contextualize each token and do things like solve co-reference and uh and and combine track where we are in a conversation um do retrieal from context and things like that. So the immediate question is when you build an audio language model that's going to do duplex conversation where what are the the tokens going to be? You might think that an an acoustic token or an audio token should represent a unit of a sound rather than a unit of a language like like text, right? So for this we would turn to audio codecs. And you might recognize MP3, AAC, Black, Wave. If you've ever streamed audio to your phone, these formats are actually audio codecs. They take raw audio waveforms and convert them into digital signals or discrete discretized um sequences of numbers like this. You have an encoder. It takes the audio in. it discretizes it that to a form that can be efficiently transmitted over a network and then the other side there's a decoder which converts it back into reconstructed waveform. Um so these dis discrete discretized uh units of audio can be our tokens and uh MP3s and a AAC is actually a little bit too um it has too high of a frame rate for us because neural networks have limited context windows. So we turn to a more efficient form of compression which is implemented as a neural audio codec and we can take um audio and compress it with into about 1.5 to 24 kilobits per second whereas MP3 is typically 96 to 320 kilobits per second. So how do how do these neural audio codecs work? Um, basically the the codec has an encoder and a decoder. It's able to take the sampled audio and output um a sequence of embedding vectors that represent the unit of sound. So at t= 1, t equals 2, all the way up to t= n, you have an embedding vector that represents in a space of possible sounds, this sound sounds similar to that sound. So we're going to embed them close together. So this is great, but this is even more information than potentially the raw audio waveform because the raw audio waveform is just a single dimensional stream of of data at a very high frequency. Whereas these vectors can can potentially be hundreds of um of uh or the dimensionality could be in the hundreds. So we use vector quantization to convert these vectors into a into integers kind of the same way you do with Vector quantization will basically take the um it it'll essentially you can think of it like k means it it segments the space up into of clusters or regions and the centrid of each of those regions represents the uh that entire region. So you have a centrid vector that represents a the typical sound that would occur in that region of the space. So you can then convert all of these units of sound vectors into yeah basically K means you can um convert these unit of sound vectors into indices which would be the indices of the code of these um centroidid vectors in a lookup table. We call that a code book and we call the individual vectors code words. So this is basically the same thing as we're doing for text except we use different terminology for it. We use uh uh embedding um embeddings and uh tokens rather than code books and code words, but it's the same thing. We now have a stream of integers and each integer represents a small unit of audio. The analogy, right? Yeah. Tokens instead of code words, embeddings instead of code uh embeddings instead of code words, tokens instead of code word indices. Um, when we have a sequence of these code words or to or audio tokens, we can then feed them back into a decoder and the decoder can reconstruct a lossy version of what was originally encoded. How lossy it is depends on the implementation. There's a whole section on something called residual vector quantization, which I actually omitted from this presentation because the sake of time, but there are methods that can be used to reduce that loss. So we have powerful neural audio codecs now that have cropped up in the last couple of years. Um these are some of the earlier ones and codec soundstream and fun codeodc. Um but at this point um I have a folder on zotterero that keeps track of all of these. It's probably tripled in size over the last year. So we have um lots of audio codecs which are basically autoenccoders with a discrete bottleneck that can convert audio to and from sequences of of of discrete tokens. So given that all we have to do is make an a an LLM except the LLM doesn't take in hex tokens. It takes in these audio tokens and we encode the input and decode the output and that's and then we're done. That's that's it. So some examples of codec language models include audiolm which was a very early one attempt at this from from Google. This was done as a three-stage language model where each stage uh would input and output tokens which represented larger and larger um units of audio. So the the first stage had the largest um they call them semantic tokens. These are large chunks of audio. Um, and the idea is that the tokens are supposed to represent the meaning of of what's being said. Whereas you have then acoustic tokens which are very granular chunks of audio and they represent um the sound properties, the acoustic properties, pitch, tone, um, intensity, all of that. So there's a separation between semantics and acoustics here. Um, audioM was a very significant step in this direction because it was the first to really try this um, and do it successfully. It actually worked both for speech and for music and um, there are some examples here uh, of what those outputs sounded like. Let's see if we can get this to play. So, I'm going to stop sharing. Um going to share again maybe directly from the Google Chrome window. Maybe that'll do it. >> This transit spring and lighting up are beautiful. >> What do you um is there is there that remote? Maybe the remote itself is muted. You look like no. I mean, >> a glamour beguiling our senses. >> This transit spring and lighting up are beautiful. >> A glamour beguiling our senses. >> Yeah. >> Kind of wants more. this transit spring and lighting up our dreams by its brilliancancy and beauty. A power which if we shut our eyes to it will not shut. >> All right, I'll play it from my laptop. I guess you guys do do the best you can to hear it. I'm sorry. Um I'm going to Or can can you mute that so it doesn't echo? >> Yeah. Which one? That one. This transit spring and lighting up are beautiful. A glamour beguiling our senses. >> This transit spring and lighting up are beautiful. Well, I think the muting the WebEx doesn't affect audio sharing. It just affects uh my voice, but whatever. >> Makes sense. >> Yeah, I was saying that because the audio is coming from the the browser, not from the the microphone. It probably whatever. I'll do it anyway just in case. >> So, um so that's the prompt. And then the continuation was sounds like this. this transit spring and lighting up our dreams by its brilliancancy and beauty. A power which if we shut our eyes to it will not shut. >> So it was able to generate speech with it the same acoustic properties although probably not very coherent but still reasonably well. Um and then it also was able to do the same thing with music. So, um, piano piano piano music. Here is the prompt. And then here is the continuation. [Music] So, so this was um I guess now I should remute myself. >> So, you can probably just continue because it's picking the recording is picking up you. >> Okay. But then we won't be able to hear anybody over the >> No, it's different. >> Okay. >> They wouldn't be able to hear the you guys. >> All right. We'll just leave it this way, I guess. Um so coming back to the slides. So yes, this was audio was a a big step in the right direction towards being able to do stuff like this. Um there was DGSLM or um which was which stood for uh spoken um dialogue language model and it took inspiration from audio LM except it only used a single um stage of codes. It didn't separate semantic and um and acoustic codes at all. But what it did is it had two different decoder heads or two actually two different decoder stacks al together. They had one for two different speakers that could be potentially talking to each other and we'll get that that's a very important detail which we'll talk about in pretty soon because that's critical for duplex language modeling. Um and I'll play some samples from that. So, this was um again a prompt and continuation type thing. So, >> hi. >> Hi, my name is D. How were you? What's your name? >> So, after the ding, it's going to play some the model's continuation of that conversation. >> Hi. >> Hi. My name is D. How were you? What's your name? >> My name is Debbie. >> Debbie, where are you calling from? Um our Carolina and it's about teen age before or right up. >> Yeah. People unabually careo. >> Oh. Well, >> yeah. But she has at least big town here from the theater one by was you can't imagine. >> Boy. Oh, Icelone. I know. And they are peace boys. They're not learning to also a peace boy. So they get the best on the theater and got a change. >> So that sounds very natural, but unfortunately the the words are not actually real, right? uh it was southern so I just I didn't understand it. >> So and and this this is a very important uh point which we'll talk about in a second. So this model had two concurrent channels one per speaker um and each channel the codes for the audio codes for each channel um went through a separate um decoder stack. They called it a dual tower LLM that had a cross attention between both stacks so that at each time step each speaker could attend to what the other speaker had said. And it used um K means um quantizer on a Hubert encoder. That's just what they used for their codec. Um and for decoding they actually used a separate it wasn't like a single autoenccoder but rather they used a separate um model for decoding. Um we also had Voli which was um a texttospech system from was it Microsoft I think right? Yeah. And um this this was a a language model which took phone and an audio enrollment sample. Basically you could do voice cloning with it. This was one of the first models that could do effective voice cloning. Now you know you go to 11 labs or you go to one of many different services and you can do very very well effective voice cloning but this was the beginning of it. Um, but all of this is the same paradigm. You have a language model. You're feeding in sequences of audio tokens and then you are generating continuations of that. Um here's another more recent one called um well they did they didn't name it this but hugging face implemented this uh reimplemented it and called it parlor or TTS but anyway um this paper introduced a texttospech system where you could feed in not only a text but also a text a a a prompt of what you would want this the speaker to sound Like so if you want the speaker to sound like they're a cowboy, you would say speak like you're a cowboy and then you would put in the the transcript and it would give you a codec decoding of that. So So we can listen to that. Maybe we can listen to it. I guess we can't listen to it. All right. >> It's not a government server, is it? Consider that. So, here's a question. Um, and I guess we probably should have heard the samples from that, but to understand the question, which is why is a LM and DGSLM unable to produce coherent speech while the texttospech models could? Um if you would have heard if that page would have loaded you would have seen that the texttospech system was very coherent. It was able to faithfully translate the text into speech. And the reason for this is that these acoustic unit embeddings they encode some information not only about what's being said but also how it's being said what what the acoustic properties are. For example the word dog spoken with in rapid speech with a rising pitch a female speaker with a southern American accent. All of that information is inside of this acoustic unit embedding. You could have many or potentially infinite embeddings for the same word. A dog spoken slowly with a level pitch um or with a falling pitch by a male speaker with a Scottish accent. Right? So same content but infinitely many possible embeddings. So if you have a speech context and you have the set of possible semantic continuations, so the words that might come next, and you have the set of possible acoustic settings, so you can think of it like a there's a cartisian product between S and A here. So for audio LM and DJS SLM, you have the space of possible continuations are in the Carteesian product of S and A. you have every possible word times every possible speaker times every possible acoustic background position. So in text to speech models like volley and the other one that I tried to show um the possible continuations are only in in um in the space of possible acoustic settings because the semantic conditioning from text constrains the at the space s. So it's a much smaller possible output space of predictions and that's why the models are able to learn coherent speech rather than just replicating the the acoustics of it. So the the interesting thing here is that audio and DJs are full language models. are actually trying to learn to reproduce speech from scratch by jointly encoding the the te you know the semantics and the acoustics in in audio tokens. The text to speech ones they're doing a much simpler modality transfer task. So if we want to the simplicity of the modality transfer task to work with full audio language modeling the solution would be interleved audio text codec language models. For example, this one speech GPT. They um take a interled sequence of text and audio and as input and they output an interleaf sequence of text and audio. Um GPT40 also does something similar to this. We don't know exactly. They didn't release the the technical specs of this, but GPT40 natively generates audio and text from the same model. Um, and its context can can contain both audio and text and it uh is able to respond within an average of uh 320 milliseconds because it does everything in a unified. Um, here is a list if you're interested. Hopefully only the people on my committee are interested um in in um of all of the interled audio text models that people have done recently. If you look at the years, I mean, this starts in 2023 and then explodes in 2024. Um, I haven't updated this for 2025. It's probably three times as long by now. So, interled audio text models have caught on and it's a a very common pattern in um in audio language models. So, this brings us to the problem we're trying to solve, which is full duplex spoken dialogue systems. So a codec language model would be full duplex if it contains a mechanism for parallel token streams. Um remember DGSLM. We'll get to that in a second. But on the left you have a half duplex codec language model where user speech comes in over a codec is converted to audio tokens. A voice activity detector will detect the end of the speech and signal the end of the turn which can be signaled end of turn token. The end of turn token then conditions the model to start generating speech which it outputs as its response. And this is very fixed. I go and turn, you go, end turn. I go, you turn, enter, you go, enter. Whereas full duplex doesn't have any of that nonsense. It just has a parallel stream of input and output codes. At every time step, you have both what's being said by on two channels by two speakers being both input and predicted. Um, and these the speech that we want to continue would be the agent speech which is passed on to the next time step. And the speech that's coming in from a microphone would would be substituted in from the codec. Of course, you could do this with both streams and, you know, generate like a selfplay type of thing. But um at inference time we would typically bring in one stream from a microphone and have the model generate the other stream. So there's no voice activity detector, no concept of an end of turn because we don't have any clear turn boundaries. And uh we have continuous parallel input and output which lends itself well to the continuous model of speech processing which I spoke about earlier. So recall DGSLM, right? They're essentially doing the same thing except rather than having one decoder stack handling both streams, they have two decoder stacks, one per stream with a cross attention between them. So it wasn't uh it's the same thing just with a different um skin on it. So more recently, there's been a couple of attempts to do full duplex dialogue modeling. Some of them worked out pretty well, such as Moshi. Moshi um is a came out of a lab called Coutai Labs which is based out of France. Um they have a they have a a two model setup. One called the helium temporal transformer which is basically modeling the first layer of a residual audio codec and they have the temporal transformer which is modeling all the other levels of the residual codec. So it it reduces the number of time steps that the um main model has to handle essentially and they use a dual stream system just like us except each stream has multiple levels. That's what I was saying. Each level models the residuals of the of the one pre below it. So here all of this is the agent speech and all of this is the user speech and each of the columns here represents a single time step of of inference. Um, see if we can hear some examples. >> Save time creating content and generate a high qual. >> Hi, thank you for joining us. We're super proud to present you Moshi today. You can see its UI actually right now on the screen. So on the top right corner, you can see some delay. We'll tell you more about it later. And you can also see some text that will be printed by the model there. So it's a multimodel model and Alex will tell you more about it later too. So let's dive straight into a demo. >> Hello, how are you today? >> Hello, can you tell me your name please? >> Hi, my name is Marshy. How can I help you today? >> Hi Moshi, can you tell me more about yourself? >> Certainly. I was created by the nonprofit research lab Qout which focuses on using AI to tackle the main challenges of modern AI. >> Okay, that sounds great. Do you know what opensource is? >> Yes, open source refers to the practice of sharing software source code free of charge. >> What are the benefits of open source? >> One of the main benefits is that it enables collaboration and allows individuals and organizations to contribute to the development of the software. >> Okay, that sounds amazing. Now, my friend Edoir has a few questions for you too if you don't mind. So that's Moshi. Um there's also >> on in that demo right at the end when he says that his colleague has some questions and she goes of course. >> Yep. That was a back channel or or actually of course could be both a back channel or and also a response. I think in that case it was a back channel. So um so this is also another recent architecture for this called sync LM. With sync LM you have um decoding set set up or segmented into chunks where each chunk contains both um audio from the user and audio from the agent. So it's more of like a flattened um a flattened interle of speakers but both um but but but still modeling both streams simultaneously because chunk n is going to have one from the user one from the uh agent and chunk n plus one is going to have one from the user one from the agent. So over time over hundreds of chunks you'll essentially have a blended um modeling of both of both streams. Um here is a list of unified and full duplex spoken dialogue systems from prior literature. These have been popping up also like really really quickly. A lot of them you most of them implement custom architectures to deal with dual stream modeling and um multi-level codecs. Um also a lot of these are are uh a lot of the sort of center of gravity for research in this area is actually in China right now. Um so the last 10 minutes we have I'll talk about the uh current my current research which is essentially I'm trying to build a full duplex codec spoken dialogive system as well as the basis of my dissertation research. Um the data sets I'm using for this are um aren't necessarily split channel. So the Fisher corpus, the call home and call friend corpora. These are old old data sets of recorded phone calls except what's important is that each speaker is isolated on a specific channel because you need each channel um each channel corresponds to each uh type of audio code. For example, the blue audio codes here are from channel one. The the gray ones are from channel two. So, if we don't have split channel audio, we can't train a model to do dual stream modeling. So, what I'm doing here is basically taking a larger subset of mono audio to pre-train the model um or to to domain adapter since I'm already starting with a pre-trained llama model and then um a small subset of split channel audio about 2,000 hours to to adapt it to for duplex dialogue. And the architecture that I'm using is actually very similar to to what I originally showed here, except instead of stacking the uh parallel channels, they're being interled in a flattened way, kind of like what Sync LM was doing. So you have one token from the user, one token from the agent, one token from the user, one token from the agent. Agent tokens are passed to the next time step. User tokens are provided from a microphone. That's it. So this allows the use of existing LLM architectures and optimized inference engines out of the box. Nothing custom here. We can use Llama 3.2 with Llama C++ immediately. Don't have to retrain it from scratch using you know trillions of tokens of text or or and millions of hours of audio. Also we're interle text and the audio to provide semantic conditioning to solve that uh incoherence problem. So we'll have interled speakers with interled text that represent what the speakers are saying. So here you have one speaker says yeah. The other speaker says that's great. And in between them you have the audio codes that represent those utterances. Um where when when B is speaking A's codes will be mostly silent with maybe some background noise. And when B A is speaking B's codes will be mostly silent also with some background noise. But they could choose at any time to overlap or to to do things like laughing together and things like that. So for a backbone, I'm using Llama 3.2 one or three billion parameters, not the instruction tuned versions because instruction tuned models, they actually enforce a half duplex turntaking mode through their chat templates, which is an interesting choice of or sort of like an assumption that's built into them that no conversations will ever be had in a duplex manner. Um, and then the codec is called Magic Codec. It is a a recent codec. It was released this year. And it basically takes 16 kHz audio um, and using a code book of 131,000 possible tokens. It can encode that 16 kHz audio into 50 tokens per channel per second. So in my model, I have 100 tokens representing a single tok second of audio from both channels. So there's four training or five training objectives. One is audio only. Just predict the next token audio token without text. And of course if this was the only training objective, we would lose coherence as we showed before. Text only so so that we don't forget how to model text, right? Um audio first. So this predicts the predict the next response token given both the audio and the text where the audio conditions the text. So audio comes first. This audio is then transcribed into this text utterance. This audio is then transcribed into this text. You can think of it as a continuous speech recognition system for it. And then the text first direction predicts the next audio token given both the audio and the text. But now the text is conditioning the audio. So you have here this text is conditioning this audio and this text is conditioning this audio. So it creates a texttospech system. And finally, the agent mode, which is what we use in real-time inference, combines those last two modes. So you have channel one uses text first or text to speech and channel 2 uses audio first or or speech recognition. So in a single stream, we have B's utterances are being transcribed from audio to text and A's utterances are being transcribed from text to text to audio. And this is an actual what an actual stream looks like. Um so we have here in blue we have the header which basically represents who are the speakers. Um and in a short enrollment of what the agent's voice should sound like. So it's like a style transfer for the agent's voice. And then you have the starting in the black section you have um interled text and audio codes where every other audio code belongs to either the user or the agent. And the what you see here on the screen is just a unic code rendering of of uh of the audio co of the codes. But there there's no meaning behind the Chinese characters or or emojis here. There's just arbitrary unic code rendering. What is it? >> It represents um each each code represents 20 milliseconds of audio from a single channel. >> Is it audio code or is it embeddings? >> These are tokens and they correspond to embeddings. >> So finally we'll get to the examples. So the in the details of the inference I'm using llama C++ as my engine. It's really nice. It's able to to run at like almost 300 tokens per second on an A100, which is great. Um, so that's about two times with all said and done, it runs about two two times real time on an A1 a single A100. Um, it runs using the agent mode with a six-second enrolled voice cloning sample and it can process audio in 100 millisecond chunks. So essentially five forward passes per second and thus can react to user input with less than 200 millisecond latency which is if you remember from the very beginning of the talk is even less than the average latency found in human dialogue. So here are some generated conversations and a real-time interactive conversation to listen to what the model sounds like. So the generated conversation the model generates both sides in audio in in text first mode and the interactive one uses agent mode. >> Hello. >> Hi. >> Hi. This is Greg. >> Hi. This is Mike. >> Hi Mike. How are you? >> Pretty good. >> Cool. So I guess we're supposed to talk about uh uh issues with the Middle East. >> Yeah. >> Okay. >> Well, I'll let you get a word in first if you >> just call the model. I didn't condition. It's a tough situation. I don't know. How do you feel about that? >> It's a tough situation. >> Yeah, I think it is too for me. >> Um um >> that uh the US needs to uh take a much more active role in the peace process. >> Yeah. >> Than we've been doing lately. >> Yeah. >> Um if we could uh if we get Israelis and Palestinian. >> Notice the the Yas that were going in there. Those are our back channels. So the other guy isn't interrupting. He's just saying, "Yeah, yeah, I'm following >> means to kind of lower the heat a little bit." I think that's the only thing that would help the situation. >> So that's what you think it's the only thing that would help it. >> Um >> the peace how do you think the peace process would work? >> Well, I think >> so one of the topics in that Fiser data set was actually the Middle East conflict. And this was back in 2004, so we know that that's been going on for a while. Um, here's another generated sample. >> I'm David. >> Hi, I'm Bob. >> How you doing? >> Not bad. How about you? >> Good. Where you calling from? >> Uh, New Jersey. >> New Jersey. Okay. I'm in San Francisco. >> Oh, well. >> So, what are we supposed to talk about? >> Um, about um Do you exercise regularly? >> Yeah. Well, I don't know. I just signed up for those things. I thought I'd better start. >> Yeah. Yeah. This is my first call. >> Oh, really? >> Hey, my third, but >> I don't do it on a regular basis. Depends where I'm at. I'm in school. >> Oh, yeah. >> It's just allowed to talk about anything, >> but not like working out or anything like that, though. >> Oh, yeah. It kind of depends on, you know, where I am at any given time. >> Oh, sure. Sure. >> Uh, so what are you into? >> Any kind of sports in particular, or are you uh related? Okay. So weight liifting and stuff like that. >> Yeah. So in the Fiser corpus participants when they collected that data participants were given one of 10 topics to talk about >> and they and then they had to talk about that for 10 minutes. So here the model that basically decides which one it's going to go with and then and then it just goes with it. >> I don't tell it, hey go talk about the Middle East. It just says well that sounds good. Next token prediction, high probability next token. Let's run with it. and it runs in it. So, um, and here is an interactive one where I'm actually interacting with with the, uh, with the model. >> Hello. >> Hello. How are you? >> Hi. Doing well. How are you? >> Good. This is David. This is Sean. How are you, David? >> Hi. Good. Um, where are you from? >> Uh, Houston. >> Nice. Um, I'm from upstate New York. >> Oh, really? My just sister and brother-in-law just moved up Houston with their daughter. >> Oh, okay. >> Where are you from? >> Oh, upstate New York. >> Oh, that's good. >> Okay. >> Yeah. Yeah. So, um what do you do for a living? >> I'm uh I do restaurant management. >> Oh, cool. You run a restaurant or do you own one or? Uh yeah, it's it's I'm managing restaurants or restaurants, but it's like I own the one that I worked for it. I do like um inventory things like just hand washing and stuff all the time to hear around there. >> Kind of quiet. >> So you're doing stuff like whatever needs to be done around the shop and you know that kind of stuff. >> Yeah. Yeah, it's a tough bone into >> um >> it's just pure generated based on during what you're >> So you essentially have the prompts. >> Yeah, in that case I I'm a live prompt. Basically my audio codes are are going in and its audio codes are being generated every uh every 100 milliseconds in interled form. >> So what's what was her name? Eliza. Eliza. >> Eliza. You don't know Liza >> like the original >> pathetic psychologist from like 1963. >> Yeah. So we'll end I know we're a little bit over time but basically there's a question of how to evaluate this and there is an existing paradigm for evaluation of conversations that you know I know LLM as a judge is used a lot now but um in 2019 there was something called acute eval which would basically would give two different conversations between two systems to a human judge and then ask a question such as which speaker would you prefer to speak to for a long conversation where speaker A came from system A and speaker B came from system B. So we can do something similar with with recorded conversations and basically ask for A to B comparisons between the the speakers and the recorded conversations and then um compute a win rate or an ELO score or something like that between different systems including one system which is actually two humans talking over the phone. Um we can also give surveys which ask participants to rate um rate A or B essentially on a bunch of different axes such as who is most most coordinated overall. Um which conversation has acknowledgements used most effectively um which has the least inappropriate interruptions etc etc. So human judges can listen and be like, "Hey, that sounds awkward, you know, or hey, this this actually sound that sounds really good and then vote for system A or system B accordingly." And if if the duplex system can win against human participants reason a reasonably well or reasonably high number of times, then that would validate the approach. So the general evaluation will be listen to a one minute conversation either selfplay or human human fill out the post survey um and compute the rate and I'll skip to that. So to conclude human speech is much richer carries much more information about than traditional text chatbot interactions and future spoken dialogue systems can be a lot more fluid and humanlike if they embrace incremental full duplex um approaches. Um and that's it. mute your WebEx. >> Wow. >> Yeah. Should be good now. >> So, yeah. The question I had was that uh St. like you were you were so how do where where is it in terms of encoding semantics into the affectations? Um so in that last example you had where you're interacting w

Original Description

Humans naturally converse in a “Full-duplex” manner, simultaneously listening, thinking, and speaking at will. Humans continuously perform ultra-low-latency turn-taking decisions that result in a conversation that is mostly coordinated but contains instances of both accidental and intentional overlap. Such overlap includes phenomena such as simultaneous laughter, sentence-completion, backchannel acknowledgements (mhm, yeah!) and interruptions. This talk will lay out the foundation for constructing naturally full-duplex AI conversational agents that replicate these phenomena, delivering a conversational experience that feels more human than AI. I will focus on language modeling techniques for interleaving the audio and text modalities, discuss recent advances in this area, and showcase the system that I am building for my dissertation research.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Playlist UU4rjm_R9sgRNvv9QsgH8LDw · Tetherless World · 37 of 40

1 TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
TWed Talk: Katie Chastain on "Breaking the Gender Schema" (6p, 24 Oct)
Tetherless World
2 TWed Talk: Neha Keshan on "Stress and Machine Learning"
TWed Talk: Neha Keshan on "Stress and Machine Learning"
Tetherless World
3 TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
TWed Talk: Sabbir Rashid on "A Semantic Data Dictionary Modelling Methods Tutorial"
Tetherless World
4 TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
TWed Talk: Brenda Thomson on "Explanation in Human-AI Systems"
Tetherless World
5 Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Spring 2019 TWed Lighting Talks: Tetherless World Constellation
Tetherless World
6 Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Twed Talk: "Global Earth Mineral Inventory: A DCO Data Legacy" (Anirudh Prabhu)
Tetherless World
7 TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
TWed Talk: Minor Gordon on "Test early, test often, and keep your master branch stable" (4 Sep 2019)
Tetherless World
8 TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
TWed Talk: Oshani Seneviratne on Ontology Aided Smart Contract Execution for Unexpected Situations
Tetherless World
9 IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
IDEA Talk: Adrien Pavao (INRIA) on Machine Learning Challenges: Crowdsourcing Big Data Problems
Tetherless World
10 TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
TWed Talk: Jim McCusker, "OWL at the Crossroads Set Theory, Graph Theory, Logic, and Computability"
Tetherless World
11 TWed Lightning Talks Fall 2019 (11 Dec 2019)
TWed Lightning Talks Fall 2019 (11 Dec 2019)
Tetherless World
12 TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
TWed Talk: Sola Shriai on "What's a Personal Health Knowledge Graph?"
Tetherless World
13 TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
TWed Talk: Minor Gordon on "A CLEAN architecture for semantic web applications" (04 Mar 2020)
Tetherless World
14 TWed Lightning Talks Spring 2020 (29 Apr 2020)
TWed Lightning Talks Spring 2020 (29 Apr 2020)
Tetherless World
15 TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
TWed Talk: Henrique Santos on "Making Sense of Common Sense" (Weds, 07 Oct 2020)
Tetherless World
16 TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
TWed Talk: Sabbir Rashid on "Annotating and Transforming Data with Semantic Data Dictionaries"
Tetherless World
17 TWed Lightning Talks (Fall 2020)
TWed Lightning Talks (Fall 2020)
Tetherless World
18 TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
TWed Talk: Sabbir Rashid on "SQuARE: The SPARQL Query Agent-based Reasoning Engine"
Tetherless World
19 TWed Lightnining Talks: Spring 2021
TWed Lightnining Talks: Spring 2021
Tetherless World
20 TWed Lightning Talks (Fall 2021)
TWed Lightning Talks (Fall 2021)
Tetherless World
21 TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
TWed Talk: Jamie McCusker on "Build Your Own Knowledge Graph With Whyis 2.0" (28 Sep 2022)
Tetherless World
22 TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
TWed Talk: Sola Shirai on "An Introduction to Rule-Learning Models for Link Prediction" 20 Oct 2022
Tetherless World
23 TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
TWed Talk (28 Feb 2023): Brenda Thomson on "Bibliometrics: The limitations and possibilities"
Tetherless World
24 TWed Lighting Talks Spring 2023
TWed Lighting Talks Spring 2023
Tetherless World
25 TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
TWed Talk (11 Oct 2023): Jamie McCusker on " "Splitting the World With My Grandfather's Axe"
Tetherless World
26 FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
FOCI LLM Users Group: "Beyond Autocomplete: Instruction Following & CoT Reasoning in LLM Agents"
Tetherless World
27 FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
FOCI GenAI Users Group (31Jan2024) : The Large Language Model for Mixed Reality (LLMR)
Tetherless World
28 TWed Lightning Talks Spring 2024 (14 Feb 2024)
TWed Lightning Talks Spring 2024 (14 Feb 2024)
Tetherless World
29 FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
FOCI LLM Users Group: "A Guide into Open Source Large Language Models and Techniques"
Tetherless World
30 Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Danielle Villa "Testing Faithfulness of Language Model-Generated Explanations" (25 Sep 2024)
Tetherless World
31 Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Jamie McCusker "Getting Started with Knowledge Graphs using Whyis" (23 Oct 2024)
Tetherless World
32 TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
TWed Talk: Tom Morgan on "Intro to Quantum Fourier Transform on the RPI Quantum One" (4p Wed 13 Nov)
Tetherless World
33 TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
TWed: Abraham Sanders on "Training Large Language Models to Reason in a Continuous Latent Space"
Tetherless World
34 TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
TWed Paper Talk: Danielle Villa on "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL"
Tetherless World
35 TWed Talk: Thilanka Munasinghe (26 Mar 2025)
TWed Talk: Thilanka Munasinghe (26 Mar 2025)
Tetherless World
36 TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
TWed Talk: "ChatBS-NexGen: A Platform for Automated KG-based LLM Fact Checking" (23 Apr 2025)
Tetherless World
"Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
"Toward Fluid AI Conversation with Natural Turn-taking: Full-duplex Modeling with Audio Codec LMs"
Tetherless World
38 TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
TWed Talk: "Detecting Ambiguity in Question Answering over Financial Documents using LLMs"
Tetherless World
39 TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
TWed Talk: "Model Context Protocol (MCP): Standardizing Tool Use for LLM Systems" (18 Feb 2026)
Tetherless World
40 TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
TWed Talk: "Discourse-Aware Scholarly Knowledge Graphs for the LLM Era" 18 Mar 2026
Tetherless World

The video teaches how to build full-duplex models with audio codec LMs for fluid AI conversation with natural turn-taking, highlighting the importance of paralinguistics and incremental speech processing. The key insight is that full-duplex modeling can be achieved by using audio codec LMs and incremental speech processing, allowing for more human-like dialogue.

Key Takeaways
  1. Pre-train LLMs with mono audio
  2. Adapt LLMs to duplex dialogue with split channel audio
  3. Interleave user and agent audio streams in a flattened way
  4. Provide semantic conditioning by interleaving text and audio
  5. Use existing LLM architectures and optimized inference engines
  6. Train models with multiple objectives, including audio only, text only, and agent mode
💡 Full-duplex modeling can be achieved by using audio codec LMs and incremental speech processing, allowing for more human-like dialogue.

Related Reads

📰
GPT-5.5 Complete Guide in 2026
Learn about GPT-5.5, its features, performance, and why it matters in 2026, and how to leverage it for improved AI-assisted tasks
Dev.to AI
📰
AI Simplified — Why Structured Output Matters More Than a Fluent Answer
Learn why structured output is crucial for AI answers and how it can improve usability and automation
Dev.to AI
📰
I ran a 110B LLM on 16GB of RAM. Here's the equation that predicts any model's speed on your machine
Learn how to predict a large language model's speed on your machine using a simple equation, and discover how to run a 110B LLM on limited RAM
Dev.to · Federico Sciuca
📰
A Fidelity-First Workflow for Editing GPT-Generated Text
Learn a fidelity-first workflow for editing GPT-generated text to improve its quality and readability
Dev.to · Bisrat
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →