arXiv:2604.16788cs.RO2026-04被引 1

构建真实世界长序列机器人操作评估基准,揭示执行与情境双重挑战。

LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks

论文配图:LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks
图 1 · 摘自论文原文
  • 基于1000+真实任务视频,分可观察与模糊情境两类评估
  • 发现长期表现受执行鲁棒性与情境复杂度共同影响
  • 适合研究机器人长期决策与抗干扰能力的团队使用

机器人操作策略在长序列任务中常出现性能退化,但现有基准难以揭示失败原因。多数以往基准基于仿真或仅报告整体成功率,无法区分真实执行中的时间性难题。本文提出LongBench,一个面向真实世界长时序操作的评估基准,包含超过1000个真实世界任务实例,涵盖两种互补范式:上下文无关(完全可观测)与上下文相关(模糊驱动)。通过将任务划分为能力与模糊性特定子集,LongBench支持对执行鲁棒性、时间一致性及情境依赖推理的机制感知评估。对六种先进策略的评测表明,长时序性能并非由单一因素决定:在完全可观测场景中,表现更依赖执行鲁棒性;而情境难度因任务而异,记忆方法并未持续提升表现。我们希望LongBench能成为研究长时序操作与开发更强鲁棒性策略的重要工具。

原文摘要 · Abstract (English)

Robotic manipulation policies often degrade over extended horizons, yet existing benchmarks provide limited insight into why such failures occur. Most prior benchmarks are either simulation-based or report aggregate success, making it difficult to disentangle the distinct sources of temporal difficulty in real-world execution. We introduce LongBench, a real-world benchmark for evaluating long-horizon manipulation. LongBench consists of over 1,000 real-world episodes, covering two complementary regimes: Context-Independent (fully observable) and Context-Dependent (ambiguity-driven). By organizing tasks into capability- and ambiguity-specific subsets, LongBench enables mechanism-aware evaluation of execution robustness, temporal consistency, and context-dependent reasoning. Evaluating six state-of-the-art policies reveals that long-horizon performance is not governed by a single factor. We observe that performance in fully observable settings is more strongly associated with execution robustness, while contextual difficulty varies across tasks and is not consistently improved by memory-based methods. We hope that LongBench serves as a useful benchmark for studying long-horizon manipulation and for developing policies with stronger robustness across both execution and contextual challenges.

机器人操作长序列任务真实世界评估鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。