探究大模型文本预训练中隐含的听觉知识及其对音频模型的影响
How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation
- 通过三种评测方式对比不同大模型的听觉知识水平
- 文本评测结果与音频任务表现高度相关,验证了文本知识可迁移性
- 为音频大模型设计提供实证依据,适合语音与多模态研究者参考
大型语言模型(LLMs)被广泛用作大型音频语言模型(LALMs)的知识骨干,但其在纯文本预训练中编码了多少听觉知识,以及这种知识如何影响下游性能仍不明确。本文通过三种设置评估不同LLM:(1) 在AKB-2000基准上直接探测听觉知识广度与深度;(2) 基于音频描述生成器输出的文本进行级联推理评估;(3) 将各LLM微调为结合音频编码器的大型音频语言模型(LALM)。结果显示,不同模型家族间的听觉知识差异显著,且纯文本评测结果与音频任务表现强相关。本工作为理解LLMs在音频研究中的作用提供了实证基础。
原文摘要 · Abstract (English)
Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains unclear. We study this gap by comparing different LLMs under two text-only and one audio-grounded setting: (1) direct probing on AKB-2000, a curated benchmark testing the breadth and depth of auditory knowledge; (2) cascade evaluation, where LLMs reason over text descriptions from an audio captioner; and (3) audio-grounded evaluation, where each LLM is fine-tuned into a Large Audio Language Model (LALM) with an audio encoder. Our findings reveal that auditory knowledge varies substantially across families, and text-only results are strongly correlated with audio performance. Our work provides empirical grounding for a comprehensive understanding of LLMs in audio research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。