提出任务状态视野评估框架,揭示智能体长时任务中的状态追踪短板。
Compiling and Benchmarking Task-State Horizons for Embodied Agents

- 构建符号化任务图谱,量化智能体需跟踪的状态演化跨度。
- 在84场景588个任务中发现多数模型在高状态视野下表现骤降。
- 适合研究长时决策、状态记忆与机器人规划的学者参考。
前沿智能体模型正被用作长时程具身任务的高层规划器。现有机器人基准主要通过动作序列长度和子任务复杂度衡量难度,却忽略了另一关键挑战:智能体必须追踪自身探索与环境动态所引发的任务相关状态演变。本文定义智能体需持续追踪的任务状态转移范围为任务状态视野(TSH)。为评估性能随TSH的变化,提出RoboGraph——一个将状态转移依赖转化为可执行符号图的任务编译器。RoboGraph基于空间与时间因果依赖构建TSH,包含执行中意外失败与外部干预引起的依赖。基于此,发布涵盖84个场景、588个任务的基准数据集,不同任务具有各异的TSH值。在语义与视觉闭环环境中对15种先进智能体模型的实验表明,多数模型在高TSH任务中表现显著下降,暴露出其在长期任务中维持、探索与更新任务相关状态能力的严重不足。
原文摘要 · Abstract (English)
Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long-horizon evaluation, but primarily characterize difficulty through action-sequence length and subtask complexity, overlooking a distinct challenge: agents must track evolving task-relevant world states induced by both their exploration and environmental dynamics. We define the span of task-relevant state transitions that an agent must track as task-state horizon (TSH). To evaluate how agent performance varies with TSH, we introduce RoboGraph, a robotic task compiler that translates state-transition dependencies into executable symbolic graphs. Specifically, RoboGraph constructs task-state horizons from spatial and temporal causal dependencies, including those induced by unexpected failures and interventions during task execution. Building on RoboGraph, we release a benchmark comprising 588 episodes across 84 scenes with varying TSHs. Experiments evaluating 15 advanced agentic models in both semantic and visual closed-loop environments show that most models struggle with demanding TSHs, revealing substantial gaps in maintaining, exploring, and updating task-relevant state over long horizon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。