arXiv:2602.11065cs.CLcs.AI2026-02

构建对话因果推理框架,实现双向语音交互的自然行为建模。

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling

  • 基于因果流式建模,分层预测对话意图与行为
  • 在真实和合成对话中实现高精度行为识别
  • 生成可解释推理链,适合交互系统研发者

人类对话由隐含的思维链条组织,并表现为时间结构化的互动行为。捕捉这一感知路径对构建自然的全双工交互系统至关重要。我们提出S-MARC(Streaming Causal Modeling and Reasoning for Conversation),一种流式、因果且分层的对话行为建模与推理框架。通过形式化意图到行为的路径,S-MARC同时预测高层沟通功能与底层交互行为,并建模其因果与时间依赖关系。为支持该设置,我们构建了一个高质量语料库,将可控、事件丰富的双工对话数据与行为标签配对。S-MARC将流式预测组织为持续演化的图结构,生成简洁决策理由,并动态优化推理过程。在合成与真实双工对话上的实验表明,S-MARC实现了稳健的行为检测,产生可解释的推理链,并为全双工语音对话系统的对话推理建立基准基础。

原文摘要 · Abstract (English)

Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing this perceptual pathway is critical for building natural full-duplex interactive systems. We propose S-MARC (Streaming Causal Modeling and Reasoning for Conversation), a streaming, causal, and hierarchical framework for conversational behavior modeling and reasoning. By formalizing the intent-to-action pathway, S-MARC predicts high-level communicative functions and low-level interaction behaviors while modeling their causal and temporal dependencies. To support this setting, we construct a high-quality corpus that pairs controllable, event-rich duplex dialogue data with behavior labels. S-MARC organizes streaming predictions into a continuously evolving graph structure, generates concise justifications for its decisions, and dynamically optimizes its reasoning process. Experiments on synthetic and real duplex dialogues show that S-MARC achieves robust behavior detection, produces interpretable reasoning chains, and establishes a benchmark foundation for conversational reasoning in full-duplex spoken dialogue systems.

对话建模因果推理全双工交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。