arXiv:2603.17837eess.AScs.CL2026-03中稿 · ICML被引 5

让对话模型在听的同时悄悄思考,提升回应质量。

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

  • 听人说话时同步进行隐式推理,不增加延迟。
  • 用变分下界目标实现高效微调,无需人工标注思考过程。
  • 适合需要自然流畅对话的实时语音系统开发者。

对话中,人类在倾听时会无意识地进行内部思考,尽管这种认知过程不总是表现为显式的语言结构,但对生成高质量回应至关重要。受此启发,我们提出一种名为FLAIR的全双工隐式内部推理方法,在语音感知过程中同步进行潜在推理。与传统NLP中需事后生成的‘思考’机制不同,该方法在用户说话阶段递归地将上一步的隐向量输入下一步,实现严格遵循因果关系的连续推理,且不引入额外延迟。为支持这一隐式推理,我们设计了一种基于证据下界(ELBO)的目标函数,可通过教师强制(teacher forcing)实现高效的监督微调,避免了对显式推理标注的需求。实验表明,这种边听边思的设计在多个语音基准上取得了有竞争力的结果,并在全双工交互指标上展现出稳健性能。

原文摘要 · Abstract (English)

During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal cognitive processing may not always manifest as explicit linguistic structures, it is instrumental in formulating high-quality responses. Inspired by this cognitive phenomenon, we propose a novel Full-duplex LAtent and Internal Reasoning method named FLAIR that conducts latent thinking simultaneously with speech perception. Unlike conventional "thinking" mechanisms in NLP, which require post-hoc generation, our approach aligns seamlessly with spoken dialogue systems: during the user's speaking phase, it recursively feeds the latent embedding output from the previous step into the next step, enabling continuous reasoning that strictly adheres to causality without introducing additional latency. To enable this latent reasoning, we design an Evidence Lower Bound-based objective that supports efficient supervised finetuning via teacher forcing, circumventing the need for explicit reasoning annotations. Experiments demonstrate the effectiveness of this think-while-listening design, which achieves competitive results on a range of speech benchmarks. Furthermore, FLAIR robustly handles conversational dynamics and attains competitive performance on full-duplex interaction metrics.

对话系统隐式推理语音生成全双工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。