arXiv:2506.15635cs.CVcs.RO2025-06被引 11

构建首个评估具身智能体长期记忆的基准,推动机器人持续决策研究。

FindingDory: A Benchmark to Evaluate Memory in Embodied Agents

  • 设计60个需长期记忆的具身任务,模拟多日交互与环境感知。
  • 现有视觉语言模型仅能处理数百张图像,难以支撑长时间推理。
  • 适合研究机器人记忆、长程规划与多模态融合的学者使用。

大型视觉语言模型在规划与控制任务中表现优异,但将其应用于真实机器人时受限于对跨日积累的长期经验(大量图像)的处理能力。当前模型通常只能并发处理几百张图像,凸显了在具身场景中高效管理长期记忆的需求。现有长视频问答基准忽视了物体操作与导航等具身挑战,这些任务需要低级技能和对过往交互的细粒度推理。有效的记忆整合需同时具备回忆历史信息与据此执行动作的能力,因此应联合考察。本文在Habitat仿真器中提出新基准,涵盖60个需持续参与与情境意识的任务,支持通过程序化扩展生成更长更复杂的版本,实现可扩展的内存与推理评估。我们还提供了结合先进视觉语言模型与底层导航策略的基线方法,评估其在这些高内存要求任务上的表现,并指明改进方向。

原文摘要 · Abstract (English)

Large vision-language models have recently demonstrated impressive performance in planning and control tasks, driving interest in their application to real-world robotics. However, deploying these models for reasoning in embodied contexts is limited by their ability to incorporate long-term experience collected across multiple days and represented by vast collections of images. Current VLMs typically struggle to process more than a few hundred images concurrently, highlighting the need for more efficient mechanisms to handle long-term memory in embodied settings. To effectively evaluate these models for long-horizon control, a benchmark must specifically target scenarios where memory is crucial for success. Existing long-video QA benchmarks overlook embodied challenges like object manipulation and navigation, which demand low-level skills and fine-grained reasoning over past interactions. Moreover, effective memory integration in embodied agents involves both recalling relevant historical information and executing actions based on that information, making it essential to study these aspects together rather than in isolation. In this work, we introduce a new benchmark for long-range embodied tasks in the Habitat simulator. This benchmark evaluates memory-based capabilities across 60 tasks requiring sustained engagement and contextual awareness in an environment. The tasks can also be procedurally extended to longer and more challenging versions, enabling scalable evaluation of memory and reasoning. We also present baselines that integrate state-of-the-art VLMs with low level navigation policies, assessing their performance on these memory-intensive tasks and highlight areas for improvement.

具身智能长期记忆机器人基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。