研究语音与文本在模型中如何记忆事实,发现记忆机制不完全通用。
Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

- 用因果中介分析法对比语音与文本的记事实现方式
- 在SpiritLM上发现语音到文本的事实召回效果低于文本到文本
- 揭示语音模态下知识存储机制不完全继承自文本模态
近年来,多种将语音与书面文本联合表示的语音语言模型(SLMs)被提出。随之而来的问题是:这些系统在两种模态下的内部机制有何异同?本文聚焦于模型如何编码、存储和检索事实知识,这一问题此前已在纯文本模型中研究过。为探究SLMs中事实关联存储与召回的内在机制,我们采用因果中介分析(Causal Mediation Analysis),该方法曾用于文本模型。初步实验基于集成离散语音标记的多模态模型SpiritLM,结果显示,从文本到文本与从语音到文本的结果存在差异,表明事实召回的涌现机制仅部分从文本模态迁移到语音模态。这些结果深化了我们对SLMs内部事实关联编码的理解,并为改进语音增强型AI系统提供了启示。
原文摘要 · Abstract (English)
In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechanisms are similar and different when operating in the two modalities. We focus on how these systems encode, store, and retrieve factual knowledge, which has previously been investigated for text-only models. To investigate mechanisms behind the storage and recall of factual association in SLMs, we leverage Causal Mediation Analysis, a technique previously applied to text-based models. Initial results using SpiritLM, a multimodal model integrating discrete speech tokens reveal discrepancies between text-to-text and speech-to-text results, suggesting that the emergent mechanisms for factual recall are only partially carried over from the text to the speech modality. These results advance our understanding of how internal mechanisms encode factual associations in SLMs while contributing insights for improving speech-enabled AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。