让机器人更精准模仿人类动作,关键在视觉语言与动作的协同对齐。
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
- 通过视觉、语言、本体感知的连续联合学习生成平滑动作轨迹。
- 在三个仿真环境中平均提升8.0%,双手插入任务最高提升19.2%。
- 适合需要高精度动作模仿的机器人控制场景,尤其复杂交互任务。
语言引导的操作通过行为克隆(BC)实现人机交互,该方法从人类示范中学习控制策略,是具身智能的核心。克服序列动作决策中的误差累积仍是提升BC性能的关键挑战。现有方法通过数据增强、表达性表征或时间抽象缓解问题,但常导致物理不连续和语义-物理错位,造成动作克隆不准、执行中断。本文提出连续视觉-语言-动作协同学习框架(CCoL),通过双向交叉注意力将语言语义锚定在视觉运动表征上,实现细粒度语义对齐,确保时序一致且精确的动作生成。实验表明,CCoL在三个仿真环境平均提升8.0%相对性能,双臂插入任务最高达19.2%提升;在7-DoF真实机器人上也验证了其对未见及噪声物体状态的泛化能力。
原文摘要 · Abstract (English)
Language-conditioned manipulation facilitates human-robot interaction via behavioral cloning (BC), which learns control policies from human demonstrations and serves as a cornerstone of embodied AI. Overcoming compounding errors in sequential action decisions remains a central challenge to improving BC performance. Existing approaches mitigate compounding errors through data augmentation, expressive representation, or temporal abstraction. However, they suffer from physical discontinuities and semantic-physical misalignment, leading to inaccurate action cloning and intermittent execution. In this paper, we present Continuous vision-language-action Co-Learning with Semantic-Physical Alignment (CCoL), a novel BC framework that ensures temporally consistent execution and fine-grained semantic grounding. It generates robust and smooth action execution trajectories through continuous co-learning across vision, language, and proprioceptive inputs (e.g., robot internal states). Meanwhile, we anchor language semantics to visuomotor representations by a bidirectional cross-attention to learn contextual information for action generation, successfully overcoming the problem of semantic-physical misalignment. Extensive experiments show that CCoL achieves an average 8.0% relative improvement across three simulation suites, with up to 19.2% relative gain in human-demonstrated bimanual insertion tasks. Real-world tests on a 7-DoF robot further confirm CCoL's generalization under unseen and noisy object states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。