研究机器人执行长任务时技能衔接失败的原因,给出可诊断的改进方向。
Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition

- 用多视角视觉语言模型验证技能切换,自动判断是否推进、重试或重规划。
- 单个技能在干净状态成功率77%~100%,但组合执行时仍频繁卡死。
- 揭示失败主因是下一技能准备就绪、目标定位和控制执行问题,适合做技能库优化参考。
长时序家庭任务需要机器人组合多个语言控制的技能,但连续技能之间的边界通常不明确。一个技能虽满足自身后置条件,却可能使机器人、物体或相机视角处于无法可靠启动下一个技能的状态。本文通过代理协调的视觉-语言-动作执行框架,在BEHAVIOR-1K数据集上研究该语义交接问题。框架调用基于π₀.₅训练的技能检查点,为每个技能分配类型化参数与步骤预算,并使用多视图视觉-语言模型验证是否推进、重试或重规划。为分离单个技能能力与长程组合鲁棒性,我们在两种初始状态分布下评估相同检查点:干净的技能边界快照与前一技能生成的链式终止状态。选定的导航、抓取、放置和开门技能在干净快照下人类审核验证成功率达77%~100%,但组合轨迹仍频繁停滞于链式状态。分析显示失败主要归因于下一技能就绪性、目标定位与控制执行问题,将近乎零的任务成功率转化为可操作的诊断,指明视觉-语言-动作技能库应学习对脏乱链式状态的鲁棒性。
原文摘要 · Abstract (English)
Long-horizon household tasks require robots to compose many language-conditioned skills, yet the boundary between consecutive skills is rarely explicit. A skill may satisfy its own postcondition while leaving the robot, objects, or camera views in a state from which the next skill cannot reliably start. We study this semantic handoff problem in BEHAVIOR-1K through an agent-orchestrated vision-language-action execution harness. The harness invokes $π_{0.5}$-based skill checkpoints trained from cleaned BEHAVIOR-1K demonstrations, assigns each skill typed arguments and a step budget, and uses multi-view vision-language model verification to decide whether execution should advance, retry, or replan. To separate isolated skill competence from long-horizon compositional robustness, we evaluate the same checkpoints under two initial-state distributions: clean skill-boundary snapshots and chained terminal states produced by previous skills. Selected navigation, grasping, placement, and door-opening skills achieve 77--100% success from clean snapshots under human-reviewed verification, yet composed rollouts still frequently stall from chained states. The resulting traces attribute failures to next-skill readiness, target grounding, and control execution, turning nearzero task success into actionable diagnostics for what VLA skill libraries must learn next: robustness to the messy chained-state distribution that clean demonstrations underrepresent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。