arXiv:2606.05737cs.CVcs.AI2026-06被引 1

让视觉语言动作模型一步生成动作,效果更好且更简单。

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

论文配图:Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 从观测到动作的单步生成比文本到图像更高效
  • 高噪声训练下在LIBERO数据集上达到95.6%成功率
  • 适合追求高效、简洁动作生成的机器人研究者

从稀疏文本生成多样化图像困难,而从丰富观测生成紧凑动作则更容易。从条件-目标视角看,视觉语言动作(VLA)模型应与图像到文本任务对齐,而非文本到图像。我们通过标准流匹配的不可约速度损失 $R_v(t,c)$ 形式化这一观点,并在受控的8模式模拟实验和MNIST图像到文本任务中验证。随后发现,高噪声训练可显著提升VLA在标准LIBERO上的单步解码性能,在LIBERO-Long上达到95.6%准确率,且在LIBERO-Plus、LIBERO-Pro及真实机器人任务中保持竞争力;而削弱条件或扩展预测时域的消融实验则可预测地消除单步优势。结果表明,单步动作生成是否有效取决于条件-目标结构,而非特殊训练策略。

原文摘要 · Abstract (English)

Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target view, Vision-Language-Action (VLA) thus aligns with image-to-text, not text-to-image. We formalize this view through the irreducible velocity loss $R_v(t,c)$ of standard flow matching and validate it with a controlled 8-mode toy experiment and image-to-text MNIST task. We then show that high-noise training boosts one-step VLA decoding on standard LIBERO, achieving 95.6% on LIBERO-Long, and remains competitive across LIBERO-Plus, LIBERO-Pro, and real-world robot tasks, while ablations that weaken the condition or expand the horizon predictably erase the one-step gain. These results suggest that whether one-step action generation works in VLA depends not on specialized training, but on the condition-target structure.

动作生成VLA模型机器人单步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。