将第一视角烹饪视频转为可执行世界,测试智能体在信息不全下的记忆与规划能力。
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

- 从真实视频中提取图转移规则,构建隐藏符号世界
- 智能体仅凭局部观察和执行反馈更新信念,任务完成率提升27%
- 强调记忆保持对复杂任务的重要性,适合评估具身智能体
居家环境中,具身智能体需在部分观测下进行规划:它们必须记住物体、追踪状态变化,并在行动失败时恢复。现有基准测试未能充分检验此能力。第一视角视频数据集虽能捕捉真实人类活动,但内容被动;交互式模拟器支持执行,却依赖合成场景和人工设计的动力学,引入仿真到现实的差距,且常假设状态完全可观测。我们提出Ego2World,一个可执行基准,将第一视角烹饪视频转化为由图转移规则驱动的可执行符号世界。基于HD-EPIC,Ego2World从视频标注中推导出可复用的转移规则,并在隐藏符号世界图中执行。评估时,模拟器维护隐藏世界图,而智能体仅依据自身局部观测与执行反馈,在信念图上进行规划。这种分离迫使智能体在未观测真实状态的情况下更新记忆并重新规划。实验表明,动作重叠分数会高估物理状态成功,而持续信念记忆能提升任务完成率并减少重复视觉探索——这表明信念维持应成为具身智能体评估的首要目标。
原文摘要 · Abstract (English)
Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egocentric video datasets capture realistic human activities but remain passive, while interactive simulators support execution but rely on synthetic scenes and hand-crafted dynamics, introducing a sim-to-real gap and often assuming fully observable state. We introduce Ego2World, an executable benchmark that turns egocentric cooking videos into executable symbolic worlds governed by graph-transition rules. Built on HD-EPIC, Ego2World derives reusable transition rules from video annotations and executes them in a hidden symbolic world graph. During evaluation, the simulator maintains the hidden world graph, while the agent plans over its own partial belief graph using only local observations and execution feedback. This separation forces agents to update memory and replan without observing the true world state. Experiments show that action-overlap scores overestimate physical-state success, and that persistent belief memory improves task completion while reducing repeated visual exploration -- suggesting that belief maintenance should be a first-class target of embodied-agent evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。