arXiv:2603.09731cs.CVcs.AI2026-03被引 1

评测大模型在第一视角下对长期动作后果的推理能力。

EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning

  • 构建新任务:从初始场景和动作序列预测最终场景。
  • 模型表现远低于人类,长时推理仍是重大挑战。
  • 适合研究具身智能与多模态大模型的学者使用。

多模态大语言模型(MLLMs)正被视为具身智能体的基础,但其能否可靠地从第一人称视角推断动作的长期物理后果仍不明确。为此,我们提出新任务——第一视角场景长期推理预测(EXPLORE-Bench):给定初始场景图像和一系列原子动作描述,模型需预测所有动作执行后的最终场景。为实现系统评估,我们基于真实第一人称视频构建了该基准,涵盖多样化场景。每个实例配以结构化最终场景标注,包括物体类别、视觉属性及物间关系,支持细粒度量化分析。在多个专有及开源的MLLM上实验表明,模型性能显著落后于人类,凸显长期第一视角推理仍是关键难题。进一步通过逐步推理进行测试时扩展分析发现,分解长动作序列可部分提升性能,但带来显著计算开销。总体而言,EXPLORE-Bench为衡量与推进第一视角具身感知的长期推理能力提供了可靠测试平台。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint. We study this gap through a new task, Egocentric Scene Prediction with LOng-horizon REasoning: given an initial-scene image and a sequence of atomic action descriptions, a model is asked to predict the final scene after all actions are executed. To enable systematic evaluation, we introduce EXPLORE-Bench, a benchmark curated from real first-person videos spanning diverse scenarios. Each instance pairs long action sequences with structured final-scene annotations, including object categories, visual attributes, and inter-object relations, which supports fine-grained, quantitative assessment. Experiments on a range of proprietary and open-source MLLMs reveal a significant performance gap to humans, indicating that long-horizon egocentric reasoning remains a major challenge. We further analyze test-time scaling via stepwise reasoning and show that decomposing long action sequences can improve performance to some extent, while incurring non-trivial computational overhead. Overall, EXPLORE-Bench provides a principled testbed for measuring and advancing long-horizon reasoning for egocentric embodied perception.

具身智能多模态长时推理第一视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。