arXiv:2512.21706cs.CLcs.AI2025-12被引 2

让语音对话系统像人一样推理对话行为,实现自然双向交流。

Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech

  • 用思维图模型建模对话意图与动作的因果关系
  • 在仿真和真实对话数据上实现高精度行为识别
  • 适合研究对话智能与可解释性交互系统的学者

人类对话由隐含的思维链驱动,表现为有时间性的言语行为。捕捉这一因果路径是构建自然全双工交互系统的关键。我们提出一个框架,通过将对话行为建模为思维图(GoT)中的因果推断,实现对对话行为的推理。该方法采用分层标注方案,预测高层次沟通意图与低层次言语行为,学习其因果与时序依赖关系。为训练系统,我们构建了一个混合语料库,包含可控制、事件丰富的仿真对话与人工标注的推理过程及真实对话语音。GoT框架将流式预测组织为动态演化的图结构,使多模态Transformer能够预测下一言语行为,生成简洁决策理由,并动态优化推理过程。在合成与真实全双工对话上的实验表明,该框架具备稳健的行为检测能力,生成可解释的推理链,并为全双工语音对话系统中的对话推理提供基准基础。

原文摘要 · Abstract (English)

Human conversation is organized by an implicit chain of thoughts that manifests as timed speech acts. Capturing this causal pathway is key to building natural full-duplex interactive systems. We introduce a framework that enables reasoning over conversational behaviors by modeling this process as causal inference within a Graph-of-Thoughts (GoT). Our approach formalizes the intent-to-action pathway with a hierarchical labeling scheme, predicting high-level communicative intents and low-level speech acts to learn their causal and temporal dependencies. To train this system, we develop a hybrid corpus that pairs controllable, event-rich simulations with human-annotated rationales and real conversational speech. The GoT framework structures streaming predictions as an evolving graph, enabling a multimodal transformer to forecast the next speech act, generate concise justifications for its decisions, and dynamically refine its reasoning. Experiments on both synthetic and real duplex dialogues show that the framework delivers robust behavior detection, produces interpretable reasoning chains, and establishes a foundation for benchmarking conversational reasoning in full duplex spoken dialogue systems.

对话推理思维图全双工语音可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。