arXiv:2606.18847cs.AI2026-06中稿 · EMNLP被引 1

构建长期记忆的家居代理评估基准,测试其跨时记忆与行动能力。

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

论文配图:WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
图 1 · 摘自论文原文
  • 设计时序扩展的家居任务追踪数据集,包含对话、动作与状态变化。
  • 在动态环境中,记忆失效与状态覆盖仍是主要挑战,计划生成困难。
  • 提出可见性感知记忆框架,适合长期交互式智能体研发者参考。

为在真实家庭中长时间辅助人类,具身智能体需记住用户习惯、世界状态及过往交互。现有长期记忆评测多聚焦语言检索与问答,而具身评测通常仅关注短时任务执行,未检验动态环境中长期记忆的实际应用。我们提出WorldLines,一个面向长期具身家庭协助的项目驱动型基准,构建包含对话、动作、执行反馈、物体与设备状态变化的时序扩展家庭轨迹,并转化为带证据链的问答与具身任务规划样本。进一步提出ObsMem,一种基于观察者的记忆框架,维护可见性感知的记忆与行动原生状态轨迹,支持状态感知决策。实验揭示在部分可观测性、状态被覆盖及长期记忆转化为具身计划方面仍存在持续挑战,而ObsMem提供了更优的参考架构。

原文摘要 · Abstract (English)

To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-centric retrieval and question answering, while embodied benchmarks often focus on short-horizon task execution without testing long-term memory use in dynamic environments. We introduce WorldLines, a project-driven benchmark for long-horizon embodied household assistance. It constructs temporally extended household traces with dialogues, actions, execution feedback, object and device state changes, and converts them into evidence-linked samples for Memory QA and Embodied Task Planning. We further propose ObsMem, an observer-grounded memory framework that maintains visibility-aware memories and action-native state trails for state-aware decisions. Experiments reveal persistent challenges in partial observability, overwritten world states, and translating long-term memory into embodied plans, while ObsMem offers a stronger reference architecture for this setting.

具身智能长期记忆家庭助手任务规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。