arXiv:2606.14084cs.RO2026-06被引 1

通过选择特定扩散噪声提升机器人动作鲁棒性与平滑度

Self-Improving VLA Policies: Selected Diffusion Noise for Spurious-Robust Action Smoothing

论文配图:Self-Improving VLA Policies: Selected Diffusion Noise for Spurious-Robust Action Smoothing
图 1 · 摘自论文原文
  • 测试时动态选择与参考噪声差异大的扩散噪声
  • 仿真中成功率提升8%,真实场景提升10%
  • 无需训练、不改模型,适合各类VLA策略

基于扩散模型的视觉-语言-动作(VLA)策略在机器人操作中具备强泛化能力,但对虚假视觉关联和动作生成噪声敏感,导致扰动下行为脆弱。本文提出一种简单、无需训练的测试时方法——选定扩散噪声(SDN),利用扩散噪声空间作为可控自由度,同时提升鲁棒性与成功率。SDN动态采样与参考集差异最大的噪声向量,减少对虚假线索的依赖;同时筛选生成更连贯动作轨迹的候选噪声。该双重目标使模型在物体遮挡观测下仍保持稳定,且降低动作抖动,无需修改模型参数。我们在两个仿真基准(Google Robot, Widow-X)和两个真实机器人数据集上评估了SDN,涵盖pi_0、Groot-N1.5、Groot-N1.6等多种VLA策略。结果表明,SDN在仿真中成功率平均提升8%,真实场景提升10%,并生成更平滑稳定的动作。研究表明,扩散噪声选择可成为提升VLA策略测试时性能的有效通用机制。

原文摘要 · Abstract (English)

Diffusion-based Vision-Language-Action (VLA) policies enable strong generalization in robotic manipulation, but remain sensitive to spurious visual correlations and noisy action generation, leading to brittle behavior under perturbations. We introduce Selected Diffusion Noise (SDN), a simple, training-free test-time method that improves both robustness and success rate by leveraging the diffusion noise space as a controllable degree of freedom. SDN dynamically samples noise vectors that are maximally separated from a reference set to mitigate reliance on spurious cues, while selecting candidates that yield more coherent action trajectories. This dual objective encourages stable behavior even under object-masked observations and reduces action jitter without modifying model parameters. We evaluate SDN on two simulation benchmarks (Google Robot, Widow-X) and two real-world robotic datasets across multiple VLA policies, including pi_0, Groot-N1.5, and Groot-N1.6. SDN consistently improves success rates by +8% in simulation and +10% in real-world settings, while producing smoother and more stable actions. Our results highlight that diffusion noise selection can serve as an effective and general mechanism for enhancing VLA policies at test time.

机器人控制扩散模型动作平滑鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。