arXiv:2604.17896cs.LGcs.AI2026-04

给视觉语言动作模型加物理可行性监督,提升机器人操作可靠性。

Can Explicit Physical Feasibility Benefit VLA Learning? An Empirical Study

论文配图:Can Explicit Physical Feasibility Benefit VLA Learning? An Empirical Study
图 1 · 摘自论文原文
  • 用几何约束设计可行性目标,显式指导模型学习物理可行动作。
  • 在低数据下训练时,可行性监督使任务成功率提升18.3%。
  • 适合关注机器人动作安全与高效学习的研究者。

视觉-语言-动作(VLA)模型通过大规模模仿学习直接将多模态输入映射为机器人动作,但现有训练方法未显式约束障碍物避让、运动学可行性等硬性物理约束。因此,物理可行行为的几何结构仅能从示范中隐含推断。本文研究显式可行性监督是否能为VLA策略提供有效结构化引导。提出一种基于几何的可行性目标,并集成至基于扩散模型的VLA策略训练中。以带障碍物的操作任务为可控探针,系统评估几何依赖的物理可行性。实验表明,引入可行性监督后,不仅提升了动作的物理可靠性,整体任务性能也显著改善,在低数据场景下学习效率提高约27%。结果表明,显式可行性信号可有效补充基于模仿的VLA学习,对构建更可靠的VLA策略具有潜力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models map multimodal inputs directly to robot actions and are typically trained through large-scale imitation learning. While this paradigm has shown strong performance, prevailing VLA training procedures do not explicitly supervise hard physical constraints such as obstacle avoidance or kinematic feasibility. As a result, the geometric structure underlying physically feasible behavior must be inferred only implicitly from demonstrations. In this paper, we study whether introducing explicit feasibility supervision can provide effective structured guidance for VLA policies. We formulate a simple geometry-grounded feasibility objective and integrate it into the training stage of a diffusion-based VLA policy. To evaluate this idea systematically, we use obstacle-aware manipulation as a controlled probe of geometry-dependent physical feasibility. Empirical results show that augmenting VLA training with feasibility supervision improves both physical reliability and overall task performance, while also enhancing learning efficiency in the low-data regime. These findings indicate that explicit feasibility signals can effectively complement imitation-based VLA learning, highlighting their potential for developing more reliable VLA policies.

机器人学习扩散模型物理约束模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。