首个针对音频感知属性的知识编辑基准,揭示大模型在听觉知识更新中的瓶颈。
SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- 构建首个面向音频感知属性的编辑基准SAKE,聚焦声学泛化能力
- 多数方法可靠但难以实现听觉泛化与多模态知识传播,序列编辑易遗忘
- 微调模态连接器优于直接修改语言模型,更具鲁棒性
知识编辑可在不重新训练的情况下实现精准更新,但以往研究集中于文本或视觉事实,忽视了抽象听觉感知知识的探索。本文提出SAKE,首个用于大型音频-语言模型(LALMs)听觉属性感知知识编辑的基准,要求修改的是声学泛化能力,而非孤立的事实。我们在三种LALMs上评估了八种不同编辑方法,在单次与序列编辑场景下考察其可靠性、泛化性、局部性和可迁移性。结果表明,多数方法虽能可靠执行编辑,但在听觉泛化、属性内局部性及跨模态知识传播方面表现不佳,且在序列编辑中常出现遗忘或退化现象。此外,微调模态连接器相比直接编辑大语言模型主干更稳健、更均衡。SAKE揭示了当前方法的关键局限,并为发展专用于听觉的LALM编辑技术奠定基础。
原文摘要 · Abstract (English)
Knowledge editing enables targeted updates without retraining, but prior work focuses on textual or visual facts, leaving abstract auditory perceptual knowledge underexplored. We introduce SAKE, the first benchmark for editing perceptual auditory attribute knowledge in large audio-language models (LALMs), which requires modifying acoustic generalization rather than isolated facts. We evaluate eight diverse editing methods on three LALMs across reliability, generality, locality, and portability, under single and sequential edits. Results show that most methods enforce edits reliably but struggle with auditory generalization, intra-attribute locality, and multimodal knowledge propagation, and often exhibit forgetting or degeneration in sequential editing. Additionally, fine-tuning the modality connector emerges as a more robust and balanced baseline compared with directly editing the LLM backbones. SAKE reveals key limitations of current methods and provides a foundation for developing auditory-specific LALM editing techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。