arXiv:2504.16936cs.MMcs.CV2025-04EMNLP被引 3

全面评估多模态大模型的音视频能力,发现其泛化强但依赖视觉、抗干扰能力优于传统模型。

Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness

  • 从有效性、效率、泛化性和鲁棒性四维度构建音视频能力评估体系
  • 零样本/少样本下表现优异,但视觉缺失或损坏时性能显著下降
  • 对对抗样本更敏感,但整体鲁棒性高于传统模型,适合音视频理解研究

多模态大语言模型(MLLMs)在处理和理解文本、音频、视觉等多元信息方面取得了显著进展。然而,目前缺乏对这类模型音视频能力的系统性评估,尤其是在分布偏移和对抗攻击等复杂场景下的表现。本文针对MLLMs的音视频能力,从有效性、效率、泛化性和鲁棒性四个关键维度展开多维度评估。通过大量实验发现,MLLMs具备强大的零样本和少样本泛化能力,可在数据有限情况下实现优异性能;但其成功高度依赖视觉模态,在视觉输入受损或缺失时性能明显下降。此外,尽管对对抗样本较为敏感,但相比传统模型展现出更强的鲁棒性。实验结果与分析为理解MLLMs的音视频能力提供了深入洞见,指出了改进方向并为未来研究提供指导。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) have recently achieved great success in processing and understanding information from diverse modalities (e.g., text, audio, and visual signals). Despite their growing popularity, there remains a lack of comprehensive evaluation measuring the audio-visual capabilities of these models, especially in diverse scenarios (e.g., distribution shifts and adversarial attacks). In this paper, we present a multifaceted evaluation of the audio-visual capability of MLLMs, focusing on four key dimensions: effectiveness, efficiency, generalizability, and robustness. Through extensive experiments, we find that MLLMs exhibit strong zero-shot and few-shot generalization abilities, enabling them to achieve great performance with limited data. However, their success relies heavily on the vision modality, which impairs performance when visual input is corrupted or missing. Additionally, while MLLMs are susceptible to adversarial samples, they demonstrate greater robustness compared to traditional models. The experimental results and our findings provide insights into the audio-visual capabilities of MLLMs, highlighting areas for improvement and offering guidance for future research.

多模态音视频理解鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。