arXiv:2505.06803cs.SDcs.CL2025-05被引 3

对比声音与视觉大模型,发现听觉感知存在短板并用跨模态蒸馏提升性能。

Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation

  • 通过人类对比实验,发现语音大模型在听觉任务中表现弱于视觉模型。
  • 跨模态蒸馏使模型在困难音源识别上准确率显著提升。
  • 适合关注多模态感知对齐与模型优化的研究者参考。

音频大语言模型(LLMs)被认为擅长识别声音对象,但其相对于其他感官模态(如视觉或视听模型)以及人类使用耳朵、眼睛或两者结合的表现尚不明确。为此,我们系统评估了音频、视觉及视听大模型(Qwen2-Audio、Qwen2-VL、Qwen2.5-Omni)在不同输入条件下(纯音频、无声视频、带音视频)识别声音对象的表现,并与人类进行对比。研究发现,Qwen2-Audio与Qwen2-VL之间存在性能差距,类似于人类耳与眼的感知差异。为缩小这一差距,我们提出一种跨模态蒸馏框架:以某一模态模型为教师,另一为学生,基于启发式模型判断音源类别对学生的挑战程度进行知识迁移。双向蒸馏(从Qwen2-VL到Qwen2-Audio,反之亦然)在困难类别上带来显著提升。该工作从人类对齐视角揭示了大模型中的感官差距,并提出了增强多模态感知能力的系统性方法。

原文摘要 · Abstract (English)

Audio large language models (LLMs) are considered experts at recognizing sound objects, yet their performance relative to LLMs in other sensory modalities, such as visual or audio-visual LLMs, and to humans using their ears, eyes, or both remains unexplored. To investigate this, we systematically evaluate audio, visual, and audio-visual LLMs, specifically Qwen2-Audio, Qwen2-VL, and Qwen2.5-Omni, against humans in recognizing sound objects of different classes from audio-only, silent video, or sounded video inputs. We uncover a performance gap between Qwen2-Audio and Qwen2-VL that parallels the sensory discrepancy between human ears and eyes. To reduce this gap, we introduce a cross-modal distillation framework, where an LLM in one modality serves as the teacher and another as the student, with knowledge transfer in sound classes predicted as more challenging to the student by a heuristic model. Distillation in both directions, from Qwen2-VL to Qwen2-Audio and vice versa, leads to notable improvements, particularly in challenging classes. This work highlights the sensory gap in LLMs from a human-aligned perspective and proposes a principled approach to enhancing modality-specific perception in multimodal LLMs.

多模态大模型感知对齐蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。