让机器人同时理解语言和视觉,自动规划复杂操作流程。
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
- 统一生成框架融合语言与视觉,用注意力机制协同建模。
- 通过动态预训练增强多模态关联,提升空间感知能力。
- 强化微调使动作逻辑与图像生成对齐,适合长程任务规划。
在复杂的具身长周期操作任务中,有效分解与执行需结合文本逻辑推理与视觉空间想象。现有方法缺乏统一的多模态规划生成框架,导致多模态规划不一致。为此,我们提出 extbf{EVLP(具身视觉-语言规划器)},一种创新的多模态统一生成框架,联合建模语言推理与视觉生成。通过包含动态预训练与强化对齐的新型训练流程,实现长周期任务的多模态规划。核心创新包括:1)统一多模态生成框架:融合语义信息与空间特征,提供全面视觉感知;生成时直接学习离散图像的联合分布,实现一步视觉合成,通过可学习的跨模态注意力机制实现语言-视觉协同建模。2)动态感知预训练:采用双向动态对齐策略,结合逆动力学与前向动力学任务,在统一特征空间中强化多模态关联。3)强化监督微调:在统一生成空间中进行基于指令的微调,构建强化损失以对齐文本动作与生成图像之间的空间逻辑,使模型具备空间感知的多模态规划能力。
原文摘要 · Abstract (English)
In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imagination to ensure efficient and accurate operation. Current methods fail to adopt a unified generation framework for multimodal planning, lead to inconsistent in multimodal planning. To address this challenge, we present \textbf{EVLP (Embodied Vision-Language Planner)}, an innovative multimodal unified generation framework that jointly models linguistic reasoning and visual generation. Our approach achieves multimodal planning for long-horizon tasks through a novel training pipeline incorporating dynamic pretraining and reinforced alignment. Our core innovations consist of three key components: \textbf{1) Unified Multimodal Generation Framework}: For understanding, We integrate semantic information with spatial features to provide comprehensive visual perception. For generation, we directly learn the joint distribution of discrete images for one-step visual synthesis, enabling coordinated language-visual modeling through learnable cross-modal attention mechanisms. \textbf{2) Dynamic Perception Pretraining}: We propose a bidirectional dynamic alignment strategy employing inverse dynamics tasks and forward dynamics tasks, effectively strengthening multimodal correlations within a unified feature space. \textbf{3) Reinforced Supervised Fine-Tuning}: While conducting instruction-based fine-tuning in the unified generation space, we construct a reinforce loss to align the spatial logic between textual actions and generated images, enabling the model to acquire spatio-awared multimodal planning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。