Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs
The authors analyze spurious onsets in full-duplex speech LLMs under prolonged silence and low input, comparing repeated sampling versus conditioning on non-speech outputs. Their experiments show onset probability spikes by over nine orders of magnitude within an 80-ms frame, supporting the hypothesis that the model’s own conditioning drives spurious speech.
They propose an inference-time causal counterfactual intervention: suppress onsets whose distributions change little if preceding user input is muted. Across 40 held-out trials per model with realistic microphone noise, the method eliminates all spurious onsets and preserves all genuine responses, with a 61 ms 95th percentile decision time, and requires no retraining.
Why it matters: Provides a real-time, retraining-free solution to reduce unwanted self-starts in full-duplex speech LLMs. Improves reliability and user experience in conversational AI systems that operate with simultaneous listening and speaking.
AI-assisted brief
official source
Add a comment
No account required