构建可解释的语音理解系统,让模型像人一样推理语音状态与动作。
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- 分模块设计语音理解流程,通过因果图连接各环节
- 在部分监督下实现可干预的反事实分析
- 适合需要透明可控语音系统的场景
当前语音-语言模型(SLMs)通常采用语音编码器与大语言模型级联结构,将语音理解视为单一黑箱。此类方法虽能较好分析语音内容,但在稀疏监督下对其他方面推理能力较弱。因此,我们主张对语音状态与行为进行显式推理,实现模块化且透明的决策。受认知科学启发,采用模块化视角与世界模型观点,使系统学习潜在状态的前向动态。我们将语音理解分解为四个通过因果图通信的模块,构建认知状态搜索空间。基于该空间的后验轨迹,指令微调的语言模型生成简洁的因果分析与面向用户的回应,支持反事实干预与部分监督下的可解释性。我们提出一种基于图的模块化语音模型,为更透明、可控的语音理解提供了新路径。
原文摘要 · Abstract (English)
Current speech-language models (SLMs) typically use a cascade of speech encoder and large language model, treating speech understanding as a single black box. They analyze the content of speech well but reason weakly about other aspects, especially under sparse supervision. Thus, we argue for explicit reasoning over speech states and actions with modular and transparent decisions. Inspired by cognitive science we adopt a modular perspective and a world model view in which the system learns forward dynamics over latent states. We factorize speech understanding into four modules that communicate through a causal graph, establishing a cognitive state search space. Guided by posterior traces from this space, an instruction-tuned language model produces a concise causal analysis and a user-facing response, enabling counterfactual interventions and interpretability under partial supervision. We present a graph-based modular speech model for explicit reasoning, highlighting a path toward more transparent and controllable speech understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。