构建首个多步跨模态的视角视频推理基准,支持时空线索分析。
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

- 引入时空密集的人工标注推理轨迹,支持复杂推理评估
- 前沿模型在提示'何处何时'后性能显著提升
- 适合研究具身智能与视频理解的科研人员
视频推理模型是视角视频与具身智能体的核心组件。现有基准仅评估输出结果(如问题答案),缺乏对中间推理步骤的评估,且多数仅限文本回答。我们提出Minerva-Ego,一个用于评估复杂视角视频理解的基准。通过扩展高质量的视角/具身场景视频数据源,添加一系列挑战性、多步骤的跨模态问题,并提供密集的时空人类标注推理轨迹。实验表明,当前最先进模型与人类表现仍有显著差距。为深入分析该差距,我们为每个推理轨迹标注了完成问题所需关注的对象及其时空掩码。大量评估显示,向前沿模型提供'何处'和'何时'的提示可显著提升性能。Minerva-Ego 可在 https://github.com/google-deepmind/neptune 下载。
原文摘要 · Abstract (English)
Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate reasoning steps, and most provide answers only in the text domain. We introduce Minerva-Ego, a benchmark for evaluating complex egocentric visual reasoning. We extend recent high-quality video data sources recorded from egocentric / embodied settings with a set of challenging, multi-step multimodal questions and spatiotemporally-dense human-annotated reasoning traces. Benchmarking experiments show that state-of-the-art models still have a large gap to human performance. To investigate this gap in detail, we annotate each reasoning trace in the dataset with the objects of interest required to solve the question, as spatiotemporal mask annotations. Through extensive evaluations, we identify that prompting frontier models with hints of 'where' and 'when' to look yields substantial improvements in performance. Minerva-Ego can be downloaded at https://github.com/google-deepmind/neptune.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。