用语言连接视觉与操控,让机器人更懂指令、更会动作。
Language-Grounded Decoupled Action Representation for Robotic Manipulation
- 将动作拆解为平移、旋转、抓取三类可解释基础单元
- 在模拟和真实场景中实现新任务零样本泛化,效果优于现有方法
- 适合需要理解自然语言指令的智能机器人研发人员
视觉语言理解与底层动作控制之间的异质性仍是机器人操作的核心挑战。尽管近期方法在特定任务上实现了动作对齐,但在新任务或语义相关任务中仍难以生成鲁棒准确的动作。为此,我们提出语言引导的解耦动作表示框架(LaDA),利用自然语言作为感知与控制间的语义桥梁。LaDA引入一个细粒度的中间层,包含三种可解释的动作基元——平移、旋转和夹爪控制,为低层动作提供明确语义结构。同时,采用语义引导的软标签对比学习目标,对齐跨任务相似动作基元,提升泛化性和运动一致性。基于课程学习思想的自适应加权策略,动态平衡对比与模仿目标,实现稳定高效的训练。在模拟基准(LIBERO 和 MimicGen)及真实世界演示中的大量实验表明,LaDA 在未见或相关任务上均表现优异,具备强泛化能力。
原文摘要 · Abstract (English)
The heterogeneity between high-level vision-language understanding and low-level action control remains a fundamental challenge in robotic manipulation. Although recent methods have advanced task-specific action alignment, they often struggle to generate robust and accurate actions for novel or semantically related tasks. To address this, we propose the Language-Grounded Decoupled Action Representation (LaDA) framework, which leverages natural language as a semantic bridge to connect perception and control. LaDA introduces a fine-grained intermediate layer of three interpretable action primitives--translation, rotation, and gripper control--providing explicit semantic structure for low-level actions. It further employs a semantic-guided soft-label contrastive learning objective to align similar action primitives across tasks, enhancing generalization and motion consistency. An adaptive weighting strategy, inspired by curriculum learning, dynamically balances contrastive and imitation objectives for stable and effective training. Extensive experiments on simulated benchmarks (LIBERO and MimicGen) and real-world demonstrations validate that LaDA achieves strong performance and generalizes effectively to unseen or related tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。