Fullduplex Signals
Weekly signals in speech-to-speech and full-duplex voice AI. Latest: 2026-W35, Aug 17 - Aug 23, 2026. Archive: fullduplex.ai/signals
Paper • 2609.03992 • PublishedNote 2026-W37 · A 3B diffusion transformer trained with flow matching on 480k hours of speech, then fine-tuned for cross-lingual dubbing, full-duplex dialogue synthesis and emotional dialogue synthesis. It runs on DAC-VAE latents mapping 48 kHz audio to 25 Hz, over 10x EnCodec's compression, and is alignment-free: alignment learned by cross-attention, no duration predictor. One-shot generation to about a minute, long-form via multi-diffusion. Reported to approach human recordings on short conversatio
Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue
Paper • 2609.03321 • PublishedNote 2026-W37 · Neural finite-state-machine dialogue puts turn-taking control tokens and response text on one causal tape under ordinary next-token prediction, keeping the base LLM's semantics at low fine-tuning cost. Its weakness has been synthetic text training data, since LLMs cannot simulate real acoustic timing. This work learns turn-taking from real human-human spoken dialogue and semantics from human-agent text, with a rule-based transformation that serialises recordings into FSM tapes without
Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
Paper • 2609.00727 • PublishedNote 2026-W37 · A mechanistic trace of speaking-style information through Whisper-large-v2, Qwen2-Audio-7B-Instruct, Qwen2.5-Omni-7B and Chroma-4B on Expresso, using centered kernel alignment, leave-one-speaker-out probes, open-ended tone prediction and a content-prosody leakage metric. All four strongly encode style in the top third of the audio encoder, and all degrade it before the output. The projector changes geometry without removing information; decoders differ in how much style survives. Mode
Perceptually Better, Semantically Worse: Measuring Speech Enhancement Impact on LLM-Based Voice Systems
Paper • 2608.30348 • PublishedNote 2026-W37 · Introduces Output Divergence Rate, the share of utterances where speech enhancement changes an LLM's intent classification relative to clean speech, benchmarked over five conditions on 2,974 SLURP clips through Whisper large-v3 and wav2vec2-large cascades. Every condition diverges significantly from zero. MetricGAN+ more than doubles ODR versus unenhanced noisy speech, 0.318 against 0.135, while improving PESQ; unmitigated echo reaches 0.836 through speaker substitution, a failure WER
microsoft/VibeVoice-ASR-Streaming-7B
Automatic Speech Recognition • 9B • Updated • 2.28k • 196Note 2026-W37 · An LLM-based end-to-end model that produces who-said-what as speech arrives, interleaving fixed-size audio chunks, a small lookahead and previous text so no separate diarization stage is needed. The technical report (arXiv 2609.02812) claims the 7B has the lowest average WER/CER across five evaluation sets and the best or tied-best speaker attribution on 12 of 13 settings. Ten languages and user-supplied hotwords. Weights for both sizes and inference code are MIT. Topped the Hugging F
Not All Fallbacks Are Failures: Understanding and Recovering from Fallbacks in Mobile Voice Assistants
Paper • 2608.30738 • PublishedNote 2026-W37 · Six months of real-world usage from more than 500 users of a smartwatch health assistant, yielding 3,030 anonymised utterances that triggered a fallback: noisy audio, transcription errors, ambiguous requests, incomplete utterances and unintended activations. The paper contributes an operational taxonomy, the annotated dataset, and a comparison of classifiers under deployment constraints, finding that lightweight embedding-based classifiers beat larger generative models on most tasks a
otoearth/otoSpeech-full-duplex-task-oriented-20h
Updated • 155 • 4Note 2026-W37 · 20.0 hours across 58 sessions and seven collaborative tasks (Spot the Difference, Photo Talk, Describe-and-Draw Portrait, Tangram Direction, Consensus Ranking, Hiring Decision, Voice-to-Form), 48 kHz channel-separated FLAC as WebDataset shards, with 11,025 timestamped interface events, per-event visibility, task stimuli and reference answers where the task has one. No transcripts. CC BY 4.0, gated with manual approval. Published 28 August, inside last week's window, and missed there.
TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue
Paper • 2608.25218 • PublishedNote 2026-W36 · A 30-hour hand-labelled corpus of dyadic conversation plus a fixed protocol for scoring end-of-turn and interruption detection, conversation type controlled across six styles, every dialogue triple-annotated at Fleiss's kappa 0.78. Fourteen systems were scored. End-of-turn recall is stable across types; interruption false positives concentrate in backchannel-dense talk. No system is simultaneously fast, high-recall and low on false positives. Disclosure: the training split is oto data
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Paper • 2608.26432 • PublishedNote 2026-W36 · Two omni-modal models converse in native audio with no external speech recognition or synthesis and no API boundary, over the unmodified tasks, tools and success checks of an established text agentic benchmark, so interaction modality is the only variable and the training loop stays local and differentiable. Existing setups either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow, or stay in text and can measure voice agents without improving them. The re
Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness
Paper • 2608.27988 • PublishedNote 2026-W36 · Models how gaze, speech and perceived interpersonal closeness signal floor changes in free four-person dialogue, using the GaMMA corpus and interpretable logistic regression over behaviourally motivated features extracted before each turn-taking event, classifying outcomes as gaps or overlaps. Most turn-taking work this window is dyadic and audio-only; this is the multi-party, multimodal case, and it deliberately trades detector accuracy for features a designer can reason about.
BreezeBlue/Breeze-TTS-2
Text-to-Speech • 3B • Updated • 8.65k • 534Note 2026-W36 · A text-to-speech release covering voice cloning, voice design and voice direction, which reached 215 likes and about 1,800 downloads in its first week and now leads the Hugging Face text-to-speech trending list. The licence is split and worth reading before integration: source code is Apache-2.0, but the model weights, any derivative models, and self-hosted outputs are restricted to research and non-commercial use. BreezeBlue is a new name in this digest and has been added to the org
nvidia/Nemotron-3-Diarization-preview
Voice Activity Detection • Updated • 512 • 48Note 2026-W36 · A streaming Sortformer speaker-diarization and speaker-tagging preview, currently the top trending voice-activity-detection repo on Hugging Face. Access is gated by manual review under the NVIDIA Software and Model Evaluation License: internal test and evaluation only, not production, only on NVIDIA GPUs, no redistribution, no using outputs or artifacts to develop another model, and no disclosure of evaluation or test results without NVIDIA's prior written consent. That last clause ma
Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
Paper • 2608.19515 • PublishedNote 2026-W35 · 480 persona-grounded scenarios hold the task fixed and vary whether the user's concern is stated in words or carried only in prosody, with objectively checkable outcomes. Giving the model the audio on top of the transcript moves the optimal-solution rate from 14.6% to 15.3%. Forcing it to first write the inferred concern into text takes the same models to 39.6%, against 40.7% for ground-truth state. The prosody is recoverable, and it still does not reach the action unless something ma
X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance
Paper • 2608.18661 • PublishedNote 2026-W35 · Synthesises from uncertain token prefixes instead of waiting for a sentence, using uncertainty-aware buffering and carrying decoder state across segment boundaries. Reported at 15.8ms median time to first token for a single request and 260.8ms at 128 concurrent. Five of its seven authors also wrote X2-Turn, the streaming turn-state model from last week, so one company is now assembling a real-time voice stack part by part without training an end-to-end duplex model at any point.
How Fragile Is Your Watermark? Training-Free Structural Removal of Neural Audio Watermarks
Paper • 2608.16566 • PublishedNote 2026-W35 · Instead of sweeping blindly through distortions, this uses cheap structural probes to locate the domain a watermark is embedded in, then applies a single attack matched to that domain, and reports a threshold-free fragility score per scheme. It needs no training and no access to the watermarking model. Anyone planning to satisfy a marking obligation with a neural audio watermark should read it before treating that watermark as the compliance artefact.
Does Listening Matter? Backchanneling and Nodding in AI Clone
Paper • 2608.19527 • PublishedNote 2026-W35 · Adds real-time predicted backchannels and head nodding to a voice-cloned avatar and measures the effect in a within-subjects study of 35 people. Perceived attentiveness, the sense of talking with the real person, and co-presence all improve significantly. The argument is that a duplex agent feels present because of how it listens, not how well it speaks, which is a case for spending latency budget on the listening side.
Towards Quantifying Benchmark Optimization in ASR Models
Paper • 2608.19936 • Published • 12Note 2026-W35 · Three behavioural probes, covering reference disagreement, masked-number recovery, and orthographic switching, show leading open ASR models reproducing verbatim benchmark reference spans even when the audio contradicts them, masks them, or leaves them ambiguous. The behaviour can be steered with a low-rank direction, which makes it a learned policy rather than an artefact. Anyone choosing a backbone off a WER leaderboard is reading a number that partly measures memorisation.
FireRedTeam/FireRedAudio
Updated • 45Note 2026-W35 · One backbone with two decoupled continuous paths, an audio encoder for understanding and a RedAE path for generation, covering ASR, audio understanding, zero-shot and instruct TTS, semantic and acoustic speech editing, and temporal grounding over recordings up to an hour. Weights are up under Apache-2.0. The benchmark claims on MMAU, MMSU, Seed-TTS-Eval and InstructTTSEval are self-reported and the linked paper is still a placeholder, so treat the scores as unverified.