新基准聚焦长时任务中的具体技能,揭示执行瓶颈。
Behavior-Skill: A Fine-Grained Benchmark for Evaluating Vision-Language-Action Policies in Long-Horizon Tasks

- 按可执行技能重构评估体系,支持独立分析
- 覆盖50个家务任务,含23万+技能实例
- 适合研究机器人长时任务失败原因的团队
长时程移动操作任务的可靠执行仍具挑战,因整体成功依赖多个子技能的完成。现有基准多基于完整任务轨迹和聚合指标,难以观测中间失败。我们提出Behavior-Skill,一个围绕可执行子技能重构学习与评估的基准。它包含10,000次示范中产生的235,492个技能实例,覆盖50个家庭任务和34种语义技能类别。每个实例配对技能指令与对齐的观测-动作片段,并关联可恢复的中间状态与技能成功条件,实现有效前条件下独立评估。我们引入轨迹级与技能级指标,超越单一任务成功率。在涵盖pi0.5和GR00T等典型视觉-语言-动作策略的完整50任务基准上实验表明,失败在技能间分布极不均匀,接触密集型操作技能构成持续瓶颈。结果证明Behavior-Skill通过暴露中间能力图谱,补充了全任务评估,助力长时程VLA策略分析与改进。数据集公开于https://github.com/nubot-nudt/Behavior-Skill。
原文摘要 · Abstract (English)
Reliable execution of long-horizon mobile manipulation tasks remains challenging because overall task success depends on the successful completion of multiple constituent skills. Existing benchmarks, however, still rely primarily on full-task rollouts and aggregate task-level metrics, making intermediate failures difficult to observe and analyze. We present Behavior-Skill, a benchmark that reformulates the learning and evaluation of long-horizon tasks around executable constituent skills. It contains 235,492 skill instances from 10,000 demonstrations across 50 household tasks and 34 semantic skill categories. Each instance pairs a skill instruction with an aligned observation-action segment, and is further associated with a restorable intermediate state and a skill success condition to enable independent evaluation under valid preconditions. We further introduce trajectory-level and skill-level metrics to characterize policy capability beyond aggregate task success. Extensive experiments across representative VLA policies including pi0.5 and GR00T on the complete 50-task benchmark show that failures are highly non-uniform across skills, with contact-rich manipulation skills forming persistent bottlenecks. These results demonstrate that Behavior-Skill complements full-task evaluation by exposing intermediate capability profiles for analyzing and improving long-horizon VLA policies. Behavior-Skill is publicly available at https://github.com/nubot-nudt/Behavior-Skill.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。