arXiv:2603.14343cs.LG2026-03中稿 · Interspeech 2026

提出首个面向语音-语言模型的知识定位与编辑框架,实现精准事实修正。

Localizing and Editing Knowledge in Large Audio-Language Models

  • 通过语音感知因果追踪定位知识存储位置
  • 音频编辑比文本编辑更有效,可实现细粒度知识更新
  • 适用于需要实时修正语音问答系统的场景

大型语音-语言模型(LALMs)在语音理解方面表现优异,使语音成为获取事实信息的自然接口。然而,这些模型基于静态语料训练,可能包含错误事实。现有模型编辑方法仅适用于纯文本大模型,未考虑连续语音表征,也未处理知识在声学或语言模块及其跨模态组件中的分布。本文构建了首个用于LALMs知识定位与编辑的音频基准,并提出一种语音驱动的定位-编辑框架。首先,利用语音感知因果追踪定位支持事实检索的层与模块,随后在识别出的位置进行编辑。实验表明,事实知识同时编码于音频与文本模块中,且音频编辑比文本编辑或微调更有效,可实现语音AI系统中细粒度的知识控制。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) have shown strong performance in speech understanding, making speech a natural interface for accessing factual information. Yet they are trained on static corpora and may encode incorrect facts. Existing model editing methods localize and update facts in text-only LLMs, but do not account for continuous speech representations, or where knowledge is stored across acoustic or language modules, or their cross-modal module. We construct the first audio benchmark for knowledge localization and editing in LALMs and propose a speech-driven locate-then-edit framework. First, we use speech-aware causal tracing to localize layers and modules that support factual retrieval and then apply editing at identified sites. Experiments show that factual knowledge is jointly encoded in audio and text modules, and that audio editing yields more effective updates than text editing or fine-tuning, enabling fine-grained knowledge control in speech AI systems.

语音生成模型编辑多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。