arXiv:2412.02611cs.CVcs.AI2024-12被引 43

测试大模型真能理解音视频信息吗?

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

  • 设计4555个音视频多模态选择题,强制模型融合视听线索
  • 闭源与开源模型普遍在音高和响度判断上表现不佳
  • 适合关注多模态理解缺陷的开发者与研究者

近期,如GPT-4o、Gemini 1.5 Pro和Reka Core等多模态大语言模型已扩展至视觉与音频模态。尽管这些模型在多种音视频应用中表现出色,但我们的DeafTest发现,它们在人类看来极其简单的任务上仍存在困难:1)判断两个声音哪个更响;2)判断两个声音哪个音调更高。基于此,我们提出AV-Odyssey Bench,一个全面的音视频基准测试,用于评估模型是否真正理解音视频信息。该基准包含4,555个精心设计的问题,每个问题均融合文本、视觉与音频成分。为确保评估精准客观,所有问题均采用多项选择形式,无需人工或大模型辅助判断。我们对一系列闭源与开源模型进行了评测,并总结观察结果。通过揭示当前模型的局限性,旨在为未来数据集构建与模型研发提供有益参考。

原文摘要 · Abstract (English)

Recently, multimodal large language models (MLLMs), such as GPT-4o, Gemini 1.5 Pro, and Reka Core, have expanded their capabilities to include vision and audio modalities. While these models demonstrate impressive performance across a wide range of audio-visual applications, our proposed DeafTest reveals that MLLMs often struggle with simple tasks humans find trivial: 1) determining which of two sounds is louder, and 2) determining which of two sounds has a higher pitch. Motivated by these observations, we introduce AV-Odyssey Bench, a comprehensive audio-visual benchmark designed to assess whether those MLLMs can truly understand the audio-visual information. This benchmark encompasses 4,555 carefully crafted problems, each incorporating text, visual, and audio components. To successfully infer answers, models must effectively leverage clues from both visual and audio inputs. To ensure precise and objective evaluation of MLLM responses, we have structured the questions as multiple-choice, eliminating the need for human evaluation or LLM-assisted assessment. We benchmark a series of closed-source and open-source models and summarize the observations. By revealing the limitations of current models, we aim to provide useful insight for future dataset collection and model development.

多模态音视频评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。