构建真实世界长序列机器人操作评估基准,揭示执行与情境双重挑战。
LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks

- 基于1000+真实任务视频,分可观察与模糊情境两类评估
- 发现长期表现受执行鲁棒性与情境复杂度共同影响
- 适合研究机器人长期决策与抗干扰能力的团队使用
机器人操作策略在长序列任务中常出现性能退化,但现有基准难以揭示失败原因。多数以往基准基于仿真或仅报告整体成功率,无法区分真实执行中的时间性难题。本文提出LongBench,一个面向真实世界长时序操作的评估基准,包含超过1000个真实世界任务实例,涵盖两种互补范式:上下文无关(完全可观测)与上下文相关(模糊驱动)。通过将任务划分为能力与模糊性特定子集,LongBench支持对执行鲁棒性、时间一致性及情境依赖推理的机制感知评估。对六种先进策略的评测表明,长时序性能并非由单一因素决定:在完全可观测场景中,表现更依赖执行鲁棒性;而情境难度因任务而异,记忆方法并未持续提升表现。我们希望LongBench能成为研究长时序操作与开发更强鲁棒性策略的重要工具。
原文摘要 · Abstract (English)
Robotic manipulation policies often degrade over extended horizons, yet existing benchmarks provide limited insight into why such failures occur. Most prior benchmarks are either simulation-based or report aggregate success, making it difficult to disentangle the distinct sources of temporal difficulty in real-world execution. We introduce LongBench, a real-world benchmark for evaluating long-horizon manipulation. LongBench consists of over 1,000 real-world episodes, covering two complementary regimes: Context-Independent (fully observable) and Context-Dependent (ambiguity-driven). By organizing tasks into capability- and ambiguity-specific subsets, LongBench enables mechanism-aware evaluation of execution robustness, temporal consistency, and context-dependent reasoning. Evaluating six state-of-the-art policies reveals that long-horizon performance is not governed by a single factor. We observe that performance in fully observable settings is more strongly associated with execution robustness, while contextual difficulty varies across tasks and is not consistently improved by memory-based methods. We hope that LongBench serves as a useful benchmark for studying long-horizon manipulation and for developing policies with stronger robustness across both execution and contextual challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。