arXiv:2512.01715cs.RO2025-12被引 6

用特征搬运成本判断视觉语言动作模型预测可靠性

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models

  • 通过计算特征搬运成本,检测动作预测的不可靠性
  • 在分布外和长序列任务中成功率提升显著,最高增益37%
  • 轻量模块可插拔,适合真实场景下的鲁棒决策

视觉-语言-动作(VLA)模型通过流匹配生成连续动作片段,但缺乏内部可靠性判断信号。分布偏移和长时程推演会使主干特征偏离动作头可可靠解码的区域,而策略无法感知或应对这种漂移。我们发现,当此类漂移发生时,观察特征向共享特征空间中的动作表示搬运成本会升高,从而提供无需额外监督的每步可靠性估计。基于此,我们提出DiG(Discrepancy Gate),一个针对流匹配VLA策略的轻量级插件模块。DiG在骨干特征与动作专家输入投影间计算切片Wasserstein搬运成本,经指数门控映射后,用于调节残差特征精炼与训练损失。推理时,该门控支持DiG-Refine,一种迭代修正动作块的流程。仿真与真实世界实验表明,DiG持续提升成功率,在分布外和长时程任务中表现最优,最大提升达37%。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models that generate continuous action chunks via flow matching lack an internal signal for judging whether a given prediction is reliable. Distribution shift and long-horizon rollouts can push backbone representations away from the region the action head decodes reliably, yet the policy has no mechanism to detect or react to this drift. We observe that the cost of transporting observation features to the action representation in a shared feature space rises precisely when such drift occurs, providing a per-step reliability estimate without extra supervision. Building on this observation, we propose DiG (Discrepancy Gate), a lightweight plug-in module for flow-matching VLA policies. DiG computes a sliced Wasserstein transport cost between backbone features and the action expert's own input projection, maps it through an exponential gate, and uses the gate to modulate both a residual feature refinement and the training loss. At inference time, the gate enables DiG-Refinefine, an iterative refinement process that corrects action chunks before execution. Experiments on both simulation and real-world scenarios show that DiG consistently improves success rates, with the largest gains under distribution shift and on long-horizon tasks.

视觉语言动作可靠性评估流匹配鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。