首个面向长时音视频理解的评测基准,填补了真实场景评估空白。
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
- 构建时长10-90分钟的长视频数据集,涵盖多模态理解任务
- 开源模型在长视频上准确率低于35%,旗舰模型最高达65%
- 适合研究长时记忆、跨模态对齐的学者和开发者
近期多模态大语言模型在音频视频理解方面取得显著进展,但现有评估仍集中于10秒至5分钟的短片段,无法反映真实应用中数十分钟长视频的需求。为此,我们提出LVOmniBench,一个专为长时音视频跨模态理解设计的新基准。该数据集包含275个来自开放平台的高质量视频,时长10至90分钟,共1,014个问答对,覆盖丰富视听动态。通过严格的人工筛选与标注,旨在全面评估模型在长期记忆、时间定位、细粒度理解及多模态感知方面的能力。实证结果表明,当前多模态大模型在处理长时输入时面临巨大挑战:开源模型平均准确率低于35%,而Gemini 3 Pro达到约65%的峰值。我们期待该数据集与发现能推动更先进模型的发展,解决长时音视频中的复杂跨模态理解问题。
原文摘要 · Abstract (English)
Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips ranging from 10 seconds to 5 minutes, failing to reflect the demands of real-world applications, where videos typically run for tens of minutes. To address this critical gap, we introduce LVOmniBench, a new benchmark designed specifically for the cross-modal comprehension of long-form audio and video. This dataset comprises high-quality videos sourced from open platforms that feature rich audio-visual dynamics. Through rigorous manual selection and annotation, LVOmniBench comprises 275 videos, ranging in duration from 10 to 90 minutes, and 1,014 question-answer (QA) pairs. LVOmniBench aims to rigorously evaluate the capabilities of OmniLLMs across domains, including long-term memory, temporal localization, fine-grained understanding, and multimodal perception. Our extensive evaluation reveals that current OmniLLMs encounter significant challenges when processing extended audio-visual inputs. Open-source models generally achieve accuracies below 35%, whereas the Gemini 3 Pro reaches a peak accuracy of approximately 65%. We anticipate that this dataset, along with our empirical findings, will stimulate further research and the development of advanced models capable of resolving complex cross-modal understanding problems within long-form audio-visual contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。