用语言指令生成物体为中心的运动流,实现无需大量实机数据的机器人操作。
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation
- 基于视觉-语言-动作模型,从图像和指令生成物体运动流。
- 在多个基准上生成流质量优于现有方法,物理实验成功率更高。
- 适合希望用少样本实现自然语言控制机器人操作的研究者。
我们提出一种基于光流的语义条件机器人操作方法,可利用人类和网络视频中的物体操作数据进行训练,仅需少量与具体机械臂相关的数据。该任务难点在于从操作前图像和自然语言指令中生成符合语义的物体轨迹,要求语言与运动流精准对齐。为此,我们提出流式视觉-语言-动作模型(LILAC),从RGB图像和自然语言指令中生成以物体为中心的2D光流,并将其转换为6-自由度机械臂轨迹。LILAC包含两个关键组件:语义对齐损失,增强语言引导以生成对齐指令的光流;提示条件跨模态适配器,将学习到的视觉提示与图像和文本特征对齐,为光流生成提供丰富线索。实验表明,本方法在多个基准上的光流生成质量优于现有方法。此外,在使用自由形式指令的物理物体操作实验中,LILAC表现出更高的任务成功率达到显著提升。项目页面见:https://lilac-75srg.kinsta.page/。
原文摘要 · Abstract (English)
We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only minimal embodiment-specific data. This task is challenging, as object trajectory generation from pre-manipulation images and natural language instructions requires appropriate instruction-flow alignment. To tackle this challenge, we propose the flow-based Language Instruction-guided open-Loop ACtion generator (LILAC). This flow-based Vision-Language-Action model (VLA) generates object-centric 2D optical flow from an RGB image and a natural language instruction, and converts the flow into a 6-DoF manipulator trajectory. LILAC incorporates two key components: Semantic Alignment Loss, which strengthens language conditioning to generate instruction-aligned optical flow, and Prompt-Conditioned Cross-Modal Adapter, which aligns learned visual prompts with image and text features to provide rich cues for flow generation. Experimentally, our method outperformed existing approaches in generated flow quality across multiple benchmarks. Furthermore, in physical object manipulation experiments using free-form instructions, LILAC demonstrated a superior task success rate compared to existing methods. The project page is available at https://lilac-75srg.kinsta.page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。