用视觉语言动作模型实现机器人隐式协作,解决提前协助问题
Learning to Assist: Collaborative VLAs for Implicit Human-Robot Collaboration

- 用端到端模仿学习训练视觉语言动作模型实现协作
- 发现动作分块策略在长任务中易因演示动作泄露导致过早协助
- 提出推理时调整方法,在不降性能前提下减少错误协助行为
人机协作结合了人类与机器的互补优势以提升任务效率。然而,现有协作系统多依赖人工设计的流水线,难以扩展至新任务。本文表明,通过模仿学习训练的端到端视觉-语言-动作(VLA)模型可支持协作操作,并分析影响其实体表现的关键因素。我们评估了两个最先进的模型,发现隐式人机协作中存在一种动作分块策略的失效模式:示范动作泄露(即动作块跨越潜在任务转换点)会导致过早协助行为。该问题随执行时长增加而加剧,真实场景中表现为机器人在人员准备前就递送工具。为此,我们提出一种推理阶段的引导方法,在保持策略性能的同时缓解错误协助。通过16名参与者参与的长时序协作装配任务实验,验证了该方法可在延长执行周期的同时减少过早协助,显著提升协作速度并降低失败率。
原文摘要 · Abstract (English)
Human-robot collaboration (HRC) combines the complementary strengths of humans and robots to improve task efficiency. However, many existing collaborative systems rely on hand-engineered pipelines, limiting their scalability and flexibility for new tasks. In this work, we show that models trained end-to-end with imitation learning, specifically vision-language-action (VLA) models, can support collaborative manipulation, and characterize the key factors affecting their real-world performance. We evaluate two state-of-the-art models and identify a failure mode of action-chunking policies in implicit HRC, where demonstration action leakage (i.e., action chunks crossing latent task transitions) can cause premature assistive behavior. We find that this issue increases with longer execution horizons and occurs in real-world collaborative VLA systems, such as when a robot attempts to hand over a tool before the person is ready. We propose an inference-time steering method to mitigate these erroneous assistive actions while preserving policy performance. Finally, through a 16-participant user study on a long-horizon collaborative assembly task, we show that steering enables a longer execution horizon while mitigating premature assistance, leading to faster collaboration and fewer failures compared to a shorter-horizon policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。