arXiv:2510.17305cs.CVcs.MM2025-10ACL被引 3

首个评估多模态模型长视频理解能力的基准,聚焦人类行为与上下文。

LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding

  • 构建1000段高信息密度长视频,融合视觉、音频、文本多模态数据。
  • 设计六类挑战任务,揭示模型在时序定位与因果推理上的不足。
  • 适合研究长视频理解、多模态融合与评测方法的学者使用。

我们提出 extbf{LongInsightBench},首个用于评估模型在长视频理解方面能力的基准,重点关注人类语言、视角、行为及其他上下文要素,并整合视觉、音频和文本三种模态。该基准在三个关键方面表现突出: extbf{a) 长时长、高信息密度视频:} 从开源数据集 FineVideo 中筛选约1,000段视频,依据时长限制及视觉与音频模态的信息密度,聚焦讲座、访谈、vlog等富含语言内容的类型。 extbf{b) 多样且具有挑战性的任务场景:} 设计六类挑战性任务,涵盖事件内(Intra-Event)与事件间(Inter-Event)任务。 extbf{c) 严谨全面的质量保障流程:} 开发三步式半自动化数据质量保障流程,确保合成问题与答案选项的难度与有效性。基于 LongInsightBench 进行系列实验,结果表明多模态模型(OLMs)在需要精确时序定位(T-Loc)和长程因果推理(CE-Caus)的任务中仍面临挑战。扩展实验揭示了多模态融合中存在的信息丢失与处理偏差。数据集与代码已公开于 https://anonymous.4open.science/r/LongInsightBench-910F/。

原文摘要 · Abstract (English)

We introduce \textbf{LongInsightBench}, the first benchmark designed to assess models' ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements, while integrating \textbf{visual, audio, and text} modalities. Our benchmark excels in three key areas: \textbf{a) Long-Duration, Information-Dense Videos:} We carefully select approximately 1,000 videos from open-source datasets FineVideo based on duration limit and the information density of both visual and audio modalities, focusing on content like lectures, interviews, and vlogs, which contain rich language elements. \textbf{b) Diverse and Challenging Task Scenarios:} We have designed six challenging task scenarios, including both Intra-Event and Inter-Event Tasks. \textbf{c) Rigorous and Comprehensive Quality Assurance Pipelines:} We have developed a three-step, semi-automated data quality assurance pipeline to ensure the difficulty and validity of the synthesized questions and answer options. Based on LongInsightBench, we designed a series of experiments. Experimental results shows that Omni-modal models(OLMs) still face challenge in tasks requiring precise temporal localization (T-Loc) and long-range causal inference (CE-Caus). Extended experiments reveal the information loss and processing bias in multi-modal fusion of OLMs. Our dataset and code is available at https://anonymous.4open.science/r/LongInsightBench-910F/.

长视频理解多模态评测因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。