新数据集EgoTempo挑战多模态模型在第一人称视频中的时间理解能力。
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
- 设计需整合全视频信息的任务,避免仅靠单帧或常识作答。
- 现有模型在新数据集上表现大幅下降,暴露时间推理短板。
- 适合研究第一人称视觉、视频时序建模的学者使用。
理解细粒度时间动态对第一人称视频至关重要,因其连续流记录频繁且近距离与物体的交互。本文指出,当前第一人称视频问答数据集常包含仅需少量帧或常识即可回答的问题,未必依赖真实视频内容。分析显示,现有先进多模态大模型在这些基准上仅用文本或单帧输入即可取得高分。为解决此问题,我们提出EgoTempo数据集,专为评估第一人称场景下的时间理解能力而设计,强调需整合整个视频信息的任务,确保模型必须依赖时间模式而非静态线索或先验知识。在EgoTempo上的广泛实验表明,当前多模态大模型在第一人称视频的时间推理方面仍显著不足。我们希望EgoTempo能推动该领域的新研究,并启发更优的时间动态建模方法。数据集与代码已公开于https://github.com/google-research-datasets/egotempo.git。
原文摘要 · Abstract (English)
Understanding fine-grained temporal dynamics is crucial in egocentric videos, where continuous streams capture frequent, close-up interactions with objects. In this work, we bring to light that current egocentric video question-answering datasets often include questions that can be answered using only few frames or commonsense reasoning, without being necessarily grounded in the actual video. Our analysis shows that state-of-the-art Multi-Modal Large Language Models (MLLMs) on these benchmarks achieve remarkably high performance using just text or a single frame as input. To address these limitations, we introduce EgoTempo, a dataset specifically designed to evaluate temporal understanding in the egocentric domain. EgoTempo emphasizes tasks that require integrating information across the entire video, ensuring that models would need to rely on temporal patterns rather than static cues or pre-existing knowledge. Extensive experiments on EgoTempo show that current MLLMs still fall short in temporal reasoning on egocentric videos, and thus we hope EgoTempo will catalyze new research in the field and inspire models that better capture the complexity of temporal dynamics. Dataset and code are available at https://github.com/google-research-datasets/egotempo.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。