仅需一次人类示范,机器人就能学会通用双臂操作。
VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- 从单次示范中分解出可复用技能,用视觉语言锚定动态调整。
- 无需重训即可应对背景变化、物体移动等干扰,支持长任务组合。
- 适合需要灵活双臂协作的工业或服务场景,跨机器人平台通用。
实现通用双臂操作需系统能高效学习极少人类输入,并适应真实世界不确定性与多样机械体。现有方法面临两难:模仿学习需大量示范以覆盖任务变体,模块化方法在动态场景中又缺乏灵活性。本文提出VLBiMan框架,通过任务感知分解,从单个真人示范中提取可复用技能,将不变的原始动作作为锚点,利用视觉语言引导动态调整可变部分。该机制在不重新训练策略的前提下,解决因背景变化、物体重置或视觉杂乱引发的场景歧义,结合语义解析与几何可行性约束。此外,系统继承人类混合控制能力,支持双臂同步与异步协同。大量实验验证了其在工具使用和多对象任务中的表现:(1)相比模仿基线,示范需求大幅降低;(2)通过原子技能拼接实现长时序任务的组合泛化;(3)对语义相似的新物体及外部扰动具有鲁棒性;(4)强跨平台迁移能力,所学技能可在不同机器人平台上直接应用而无需重训。本工作通过融合人类先验与视觉语言锚定自适应,推动非结构化环境中实用且多功能的双臂操作发展。
原文摘要 · Abstract (English)
Achieving generalizable bimanual manipulation requires systems that can learn efficiently from minimal human input while adapting to real-world uncertainties and diverse embodiments. Existing approaches face a dilemma: imitation policy learning demands extensive demonstrations to cover task variations, while modular methods often lack flexibility in dynamic scenes. We introduce VLBiMan, a framework that derives reusable skills from a single human example through task-aware decomposition, preserving invariant primitives as anchors while dynamically adapting adjustable components via vision-language grounding. This adaptation mechanism resolves scene ambiguities caused by background changes, object repositioning, or visual clutter without policy retraining, leveraging semantic parsing and geometric feasibility constraints. Moreover, the system inherits human-like hybrid control capabilities, enabling mixed synchronous and asynchronous use of both arms. Extensive experiments validate VLBiMan across tool-use and multi-object tasks, demonstrating: (1) a drastic reduction in demonstration requirements compared to imitation baselines, (2) compositional generalization through atomic skill splicing for long-horizon tasks, (3) robustness to novel but semantically similar objects and external disturbances, and (4) strong cross-embodiment transfer, showing that skills learned from human demonstrations can be instantiated on different robotic platforms without retraining. By bridging human priors with vision-language anchored adaptation, our work takes a step toward practical and versatile dual-arm manipulation in unstructured settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。