首个评估音视频大模型可靠性的基准,发现多数模型表现远不如人类。
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
- 构建60万样本的音视频可信度评测基准,覆盖对抗攻击等三维度。
- 13个主流模型在9项任务中平均性能仅达人类水平的69.81%。
- 提出通用优化方法CAVPref,性能最高提升30.19%,适合模型鲁棒性研究者。
随着多模态大语言模型(MLLMs)的快速发展,现有诊断基准主要聚焦视觉能力,缺乏对音视频(AV)整体理解的评估。此外,尚无基准考察模型在输入扰动下的响应校准能力。为此,我们提出音视频可信度评估基准(AVTrustBench),包含60万样本、跨越9个精心设计的任务,从对抗攻击、组合推理和模态依赖三个维度评估音视频大模型(AVLLMs)。通过该基准,我们系统评估了13个前沿AVLLMs,结果表明多数模型在综合理解上显著落后于人类水平,为未来研究提供重要洞见。为进一步缓解现有方法局限,我们提出一种模型无关的校准音视频偏好优化训练策略CAVPref,跨9项任务性能最高提升30.19%。代码与数据集将公开发布,以推动该领域发展。
原文摘要 · Abstract (English)
With the rapid advancement of Multi-modal Large Language Models (MLLMs), several diagnostic benchmarks have recently been developed to assess these models' multi-modal reasoning proficiency. However, these benchmarks are restricted to assessing primarily the visual aspect and do not examine the holistic audio-visual (AV) understanding. Moreover, currently, there are no benchmarks that investigate the capabilities of AVLLMs to calibrate their responses when presented with perturbed inputs. To this end, we introduce Audio-Visual Trustworthiness assessment Benchmark (AVTrustBench), comprising 600K samples spanning over 9 meticulously crafted tasks, evaluating the capabilities of AVLLMs across three distinct dimensions: Adversarial attack, Compositional reasoning, and Modality-specific dependency. Using our benchmark we extensively evaluate 13 state-of-the-art AVLLMs. The findings reveal that the majority of existing models fall significantly short of achieving human-like comprehension, offering valuable insights for future research directions. To alleviate the limitations in the existing approaches, we further propose a robust, model-agnostic calibrated audio-visual preference optimization based training strategy CAVPref, obtaining a gain up to 30.19% across all 9 tasks. We will publicly release our code and benchmark to facilitate future research in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。