arXiv:2502.15679cs.ROcs.AI2025-02被引 7

发现视觉伺服机器人长任务失败主因:观察空间漂移,提出新基准测试

BOSS: Benchmark for Observation Space Shift in Long-Horizon Task

  • 构建三种挑战场景,量化评估技能链中观察空间漂移问题
  • 多算法测试显示性能平均下降34%至67%,证明漂移严重影响表现
  • 适合研究长时序任务、模仿学习与机器人视觉伺服的学者参考

机器人长期任务中,即使采用分层结构执行技能组合,仍常因观察空间漂移(OSS)导致性能下降。该现象指前序技能执行引发后续技能观察分布偏移,破坏预训练策略有效性。为此,我们提出BOSS基准,包含三个挑战:单谓词漂移、累积谓词漂移和技能链任务,用于系统评估此问题。在多个主流模仿学习算法上测试,包括三种行为克隆方法和OpenVLA模型,在最简单挑战下性能平均下降34%至67%。此外,增加训练数据多样性虽提升鲁棒性,但仍不足以解决根本问题。项目主页:https://boss-benchmark.github.io/

原文摘要 · Abstract (English)

Robotics has long sought to develop visual-servoing robots capable of completing previously unseen long-horizon tasks. Hierarchical approaches offer a pathway for achieving this goal by executing skill combinations arranged by a task planner, with each visuomotor skill pre-trained using a specific imitation learning (IL) algorithm. However, even in simple long-horizon tasks like skill chaining, hierarchical approaches often struggle due to a problem we identify as Observation Space Shift (OSS), where the sequential execution of preceding skills causes shifts in the observation space, disrupting the performance of subsequent individually trained skill policies. To validate OSS and evaluate its impact on long-horizon tasks, we introduce BOSS (a Benchmark for Observation Space Shift). BOSS comprises three distinct challenges: "Single Predicate Shift", "Accumulated Predicate Shift", and "Skill Chaining", each designed to assess a different aspect of OSS's negative effect. We evaluated several recent popular IL algorithms on BOSS, including three Behavioral Cloning methods and the Visual Language Action model OpenVLA. Even on the simplest challenge, we observed average performance drops of 67%, 35%, 34%, and 54%, respectively, when comparing skill performance with and without OSS. Additionally, we investigate a potential solution to OSS that scales up the training data for each skill with a larger and more visually diverse set of demonstrations, with our results showing it is not sufficient to resolve OSS. The project page is: https://boss-benchmark.github.io/

机器人模仿学习长时序任务视觉伺服

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。