arXiv:2502.14145cs.CLeess.AS2025-02被引 19

用轻量大模型实现全双工对话的实时换轮管理

LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

  • 用0.5B参数微调大模型做语义语音活动检测,判断何时该说话
  • 能区分有意打断和无意干扰,准确识别用户停顿与提问结束
  • 可独立优化对话管理模块,不需重训练核心引擎

全双工语音对话系统需实时协调听、说、思。本文提出一种语义语音活动检测(VAD)模块作为对话管理器(DM),高效管理换轮。该模块为在全双工对话数据上微调的轻量级(0.5B)大模型,可预测四个控制标记以调节换轮与持续发言,区分有意/无意打断,并识别查询完成以处理用户停顿与犹豫。通过短时间隔处理输入语音,实现实时决策;核心对话引擎(CDE)仅在生成回应时激活,降低计算开销。该设计支持对话管理模块独立优化,无需重训CDE,兼顾交互准确率与推理效率,适用于可扩展的下一代全双工对话系统。

原文摘要 · Abstract (English)

Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This paper proposes a semantic voice activity detection (VAD) module as a dialogue manager (DM) to efficiently manage turn-taking in full-duplex SDS. Implemented as a lightweight (0.5B) LLM fine-tuned on full-duplex conversation data, the semantic VAD predicts four control tokens to regulate turn-switching and turn-keeping, distinguishing between intentional and unintentional barge-ins while detecting query completion for handling user pauses and hesitations. By processing input speech in short intervals, the semantic VAD enables real-time decision-making, while the core dialogue engine (CDE) is only activated for response generation, reducing computational overhead. This design allows independent DM optimization without retraining the CDE, balancing interaction accuracy and inference efficiency for scalable, next-generation full-duplex SDS.

全双工对话对话管理大模型应用语音检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。