arXiv:2608.28630cs.CLcs.AI2026-09

让对话系统像人一样实时打断和回应,更自然流畅。

Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework

论文配图:Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework
图 1 · 摘自论文原文
  • 提出轻量级实时控制模块,可无缝接入现有模型实现主动对话
  • 构建2981小时真实对话数据集,支持五种互动风格精细标注
  • 创新双层评估体系,兼顾时机精准与交互质量,适合语音助手研发

相比需等待用户说完才回应的半双工对话系统,自然的全双工系统要求智能体能实时主动介入,包括适时打断和回应。这带来核心挑战:提升发言时机准确性的同时不牺牲回应质量。为解决现实主动对话的局限性,本文构建了一个通用风格感知的全双工框架,包含三大组件:首先,提出LPS-TC——一个即插即用的轻量级主动发言控制模块,其细粒度动作空间涵盖反应式与主动性行为,可为半双工模型赋予全双工能力,并提升已有全双工模型的时机控制能力;其次,构建了WildTurn数据集,包含约2981小时经过筛选的真实世界英语多轮立体对话(来自面对面与电话场景),并标注了五种话轮转换与五种回应风格;在该数据集上训练的LPS-TC展现出现有静态全双工基准无法捕捉的丰富口语动态;第三,提出两层评估方案,评估在真实流式约束下的块级时机精度与话轮级交互质量。实验中,将LPS-TC集成至半双工模型Qwen2.5-Omni及全双工模型Freeze-Omni,均表现出更优的时机恰当性与回应质量。该框架还展示出细粒度风格可控性与强泛化能力,推动更自然、类人化的口语交互发展。

原文摘要 · Abstract (English)

Compared with half-duplex dialogue systems where the system waits for user turn completion before it responds, natural full-duplex dialogue systems require agents to act proactively in real time, including timely interruptions and backchannels. This creates a key challenge: improving turn timing without sacrificing response quality. To address limitations in realistic proactive turn-taking, we build a generalized style-aware full-duplex framework with three key components. Firstly, we propose LPS-TC, a Lightweight Proactive Speech Turn Controller for plug-and-play integration. It features a fine-grained action space covering both reactive and proactive turn behaviors, enabling half-duplex models with full-duplex capabilities and enhancing existing full-duplex models with superior timing control. Secondly, we construct WildTurn, a large-scale, real-world English dataset containing approximately 2,981 hours of filtered multi-turn stereo conversations from face-to-face and telephone conversations, annotated with five turn-taking and five backchanneling styles. Trained on WildTurn, LPS-TC exhibits rich spoken dynamics that are not captured by existing static full-duplex benchmarks. Thirdly, we introduce a two-tier evaluation scheme that assesses both chunk-level timing precision and turn-level interaction quality under realistic streaming constraints. Our experiments, integrating LPS-TC with half-duplex models like Qwen2.5-Omni and full-duplex models like Freeze-Omni, showcase its superior performance in timing appropriateness and response quality. Our framework also demonstrates fine-grained style controllability and strong generalizability, enabling more natural and human-like spoken interactions.

语音交互全双工对话系统风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。