构建首个面向长视频多模态推理的可解释评估基准,支持跨模态、长时序、开放式问答。
A Benchmark for Omni-Modal Reasoning in Long Videos
- 设计覆盖视觉、语音、环境音的多模态融合评估框架,支持多轮开放问题
- 引入评分细粒度诊断体系,可定位模型在感知、时序、推理等环节的缺失
- 提出无需训练的证据搜索代理,实现多模态证据迭代验证,适合研究长视频理解
长视频的全模态理解需要融合视觉、语音与环境音频,并具备连贯的长上下文推理能力。现有视频评测基准常在时间尺度、模态覆盖、开放交互和可解释评分之间权衡。为此,我们提出LongShOTBench,一个围绕三大目标设计的长视频理解评测基准:全模态整合、意图驱动的开放式交互、以及评分级别的诊断分析。基准基于真实观看场景构建单轮与多轮问题,系统性地测试视觉、语音、环境音频、时间关系及跨模态推理能力。每个题目包含参考答案与加权的细粒度评分标准,可识别模型在感知事实、时间关联、模态对齐与推理步骤上的满足或遗漏情况。所有样本均经人工校验,以提升语义对齐性、清晰度与评分可靠性。我们还引入LongShOTAgent,一种无需训练的全模态证据搜寻代理,结合全视频预处理、目标检索、查询自适应片段精炼与视觉、语音及非语音音频证据的显式命题验证。其迭代的搜索-精炼-验证循环可暴露中间证据,允许特定模态专家重新分析相关时刻后作答。我们评估了105个视频能力模型,涵盖开源多模态大模型、视觉语言系统、音频大模型、智能体流水线与闭源API。当前多模态大模型仍远未达到基准上限,而LongShOTAgent是表现最强的免训练系统,整体得分达66.64%。通过发布基准、排行榜与方法代码,我们提供了一个共享且可解释的长视频多模态推理评测平台。代码、数据与排行榜详见https://longshot.cvmbzuai.com/。
原文摘要 · Abstract (English)
Long-form omni-modal video understanding requires integrating vision, speech, and ambient audio with coherent long-context reasoning. Existing video benchmarks often trade off temporal scale, modality coverage, open-ended interaction, and interpretable scoring. To address this gap, we introduce LongShOTBench, a long video understanding benchmark designed around three coupled goals: holistic omni-modal integration, intent-driven open-ended interaction, and rubric-level diagnosis. It builds single- and multi-turn questions from real viewing scenarios, with systematic tasks probing visual, speech, ambient-audio, temporal, and cross-modal reasoning. Each item includes a reference answer and a weighted criterion-level rubric, letting evaluation identify which perceptual facts, temporal links, modality-grounding requirements, and reasoning steps are satisfied or missed. All samples are manually verified to improve grounding, clarity, and rubric reliability. We also introduce LongShOTAgent, a training-free omni-modal evidence-seeking agent coupling full-video preprocessing with targeted retrieval, query-adaptive segment refinement, and explicit claim verification over visual, speech, and non-speech audio evidence. Its iterative search-refine-verify loop exposes intermediate evidence and lets modality-specific specialists re-analyze relevant moments before answering. We evaluate 105 video-capable models spanning open-source omni-modal models, vision-language systems, audio LLMs, agentic pipelines and closed-source APIs. Current MLLMs remain far from saturating LongShOTBench, while our LongShOTAgent is the strongest training-free system, reaching 66.64% overall. By releasing the benchmark, leaderboard, and method, we provide a shared, interpretable testbed for advancing long-form omni-modal video reasoning. Code, data, and the leaderboard are available at https://longshot.cvmbzuai.com/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。