让视觉语言动作模型学会用3D结构蓝图指导精准操作。
Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models

- 用3D信息构建可执行的物体程序化蓝图
- 提升复杂操作的成功率和学习效率
- 适合需要高精度物理交互的机器人研究
当前视觉语言动作(VLA)模型主要依赖2D输入,忽略了三维世界中物体的结构信息与常识知识,限制了其空间感知与复杂高精度操作的适应能力。为填补这一关键差距,我们构建了概念专家模块,使VLA能够生成代表物体的显式、可执行的解析概念。该机制分两阶段协同工作:首先,在VLA推理前,概念专家利用视觉基础模型(VFMs)的3D信息估计初始运动学与结构参数;其次,在操作过程中,VLA模型动态追踪概念参数变化,持续对齐观测结果以保持准确性。建立后的解析概念通过密集的程序化操作奖励和精确的空间引导,为VLA微调提供高质量指导,使其在保持端到端学习灵活性的同时,学会基于物理的交互行为。实验表明,在监督与强化学习场景下,成功率与学习效率均有稳定提升,验证了基于结构化概念引导在VLA后训练中的有效性。
原文摘要 · Abstract (English)
Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inherent in the 3D physical world. This deficiency restricts their spatial awareness and adaptability for complex, high-precision manipulation. To bridge this crucial gap, we construct a Concept Expert module for VLA to build executable Analytic Concepts that represent objects as explicit, programmatic blueprints. Our mechanism operates in two synergistic phases: First, prior to VLA inference, the Concept Expert leverages 3D information from Vision Foundation Models (VFMs) to estimate the initial kinematic and structural parameters. Second, throughout the manipulation process, the VLA model utilizes its inherent capability to dynamically track the dynamic concept parameters, continuously aligning them with observational changes to ensure persistent accuracy. Once established, the Analytic Concepts provide explicit, high-quality guidance for VLA fine-tuning through (1) dense, programmatic manipulation rewards and (2) precise spatial guidance. This formulation allows VLA models to learn physically grounded interaction behaviors while maintaining end-to-end learning flexibility. Our experimental results show consistent improvements in success rate and learning efficiency across supervised and reinforcement learning settings, demonstrating the effectiveness of structured, concept-based guidance for VLA post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。