构建复杂视频推理新基准,考验多步视觉理解与逻辑推理能力
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
- 需整合多个时间片段的视觉证据,结合逻辑推理回答问题
- 人类准确率仅18.97%(禁重看),顶尖模型最高45.96%
- 适合研究长时程感知推理与多模态大模型的学者使用
我们提出PerceptionComp,一个手动标注的复杂、长时程、以感知为中心的视频推理基准。该基准要求每个问题必须依赖多个时间上分离的视觉线索,并在合取与序列逻辑下进行组合约束,涵盖物体、属性、关系、位置、动作、事件等感知子任务,需要语义识别、视觉对应、时间推理和空间推理等技能。数据集包含279段来自城市漫步、别墅室内导览、游戏视频和极限户外运动等多样领域的视频,共1,114个高度复杂的问答对,100%人工标注。人类实验表明,PerceptionComp需大量测试时思考和重复感知:参与者耗时显著更长,禁用重看时准确率降至近随机水平(18.97%)。现有最先进多模态大模型表现也大幅下降:评估中最佳模型Gemini-3-Flash在五选一设置下准确率仅为45.96%,开源模型普遍低于40%。结果表明,以感知为核心的长时程视频推理仍是重大瓶颈,我们期望PerceptionComp能推动该领域进展。
原文摘要 · Abstract (English)
We introduce PerceptionComp, a manually annotated benchmark for complex, long-horizon, perception-centric video reasoning. PerceptionComp is designed so that no single moment is sufficient: answering each question requires multiple temporally separated pieces of visual evidence and compositional constraints under conjunctive and sequential logic, spanning perceptual subtasks such as objects, attributes, relations, locations, actions, and events, and requiring skills including semantic recognition, visual correspondence, temporal reasoning, and spatial reasoning. The benchmark contains 1,114 highly complex questions on 279 videos from diverse domains including city walk tours, indoor villa tours, video games, and extreme outdoor sports, with 100% manual annotation. Human studies show that PerceptionComp requires substantial test-time thinking and repeated perception steps: participants take much longer than on prior benchmarks, and accuracy drops to near chance (18.97%) when rewatching is disallowed. State-of-the-art MLLMs also perform substantially worse on PerceptionComp than on existing benchmarks: the best model in our evaluation, Gemini-3-Flash, reaches only 45.96% accuracy in the five-choice setting, while open-source models remain below 40%. These results suggest that perception-centric long-horizon video reasoning remains a major bottleneck, and we hope PerceptionComp will help drive progress in perceptual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。