arXiv:2510.02066cs.CLcs.SD2025-10被引 5

提出流式思维链框架,让语音对话系统更连贯、低延迟地实时交互。

Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems

  • 分块处理输入与生成,用帧级对齐构建中间目标
  • 响应更连贯可解释,延迟更低,支持重叠对话
  • 适合需要实时交互的智能客服、助手场景

多数端到端语音对话系统依赖语音活动检测(VAD)判断发言轮次,但VAD无法区分停顿与语句结束。全双工系统通过连续输出(含静音标记)解决此问题,但常采用复杂双通道结构,且语义推理能力落后于串行模型。为此,我们提出SCoT:一种面向全双工系统的流式思维链框架,交替处理固定时长用户输入与分块生成响应。利用帧级对齐,为每个块创建对齐的用户转录与系统响应作为中间目标。实验表明,该方法生成的响应更具连贯性与可解释性,同时支持更低延迟和重叠交互,优于现有全双工方法,并在实时性上超越传统轮次制系统。

原文摘要 · Abstract (English)

Most end-to-end (E2E) spoken dialogue systems (SDS) rely on voice activity detection (VAD) for turn-taking, but VAD fails to distinguish between pauses and turn completions. Duplex SDS models address this by predicting output continuously, including silence tokens, thus removing the need for explicit VAD. However, they often have complex dual-channel architecture and lag behind cascaded models in semantic reasoning. To overcome these challenges, we propose SCoT: a Streaming Chain-of-Thought (CoT) framework for Duplex SDS, alternating between processing fixed-duration user input and generating responses in a blockwise manner. Using frame-level alignments, we create intermediate targets-aligned user transcripts and system responses for each block. Experiments show that our approach produces more coherent and interpretable responses than existing duplex methods while supporting lower-latency and overlapping interactions compared to turn-by-turn systems.

语音对话流式处理思维链全双工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。