arXiv:2602.01158cs.CVcs.RO2026-02被引 2

修复视觉污染可显著提升机器人模型在真实场景的鲁棒性

Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs

  • 设计通用视觉修复模型CRT,实时恢复被干扰的图像
  • 在真实数据上使先进模型成功率从2%回升至90%
  • 无需微调主模型,适合部署于各类视觉-语言-动作系统

视觉-语言-动作(VLA)模型已成为通用机器人操作的主流范式,将感知与控制统一于端到端架构中。然而,尽管在受控环境中表现优异,其在真实世界中的可靠部署严重受限于对视觉干扰的脆弱性。现有研究多关注由场景几何导致的物理遮挡,而对传感器层面的图像污染——如电子噪声、坏点、镜片污渍等——却鲜有探讨。本文量化了该问题的严重性:最先进的VLAs(如$π_{0.5}$和SmolVLA)在常见信号缺陷下,成功率从90%骤降至最低2%。为此,我们提出腐蚀恢复变压器(CRT),一种即插即用、模型无关的视觉变换器,通过对抗训练目标,在不需对主模型进行昂贵微调的情况下,从受损输入中恢复清晰观测。在LIBERO和Meta-World基准上的大量实验表明,CRT能有效恢复性能,使VLAs在严重视觉污染下仍保持接近基线的成功率。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a dominant paradigm for generalist robotic manipulation, unifying perception and control within a single end-to-end architecture. However, despite their success in controlled environments, reliable real-world deployment is severely hindered by their fragility to visual disturbances. While existing literature extensively addresses physical occlusions caused by scene geometry, a critical mode remains largely unexplored: image corruptions. These sensor-level artifacts, ranging from electronic noise and dead pixels to lens contaminants, directly compromise the integrity of the visual signal prior to interpretation. In this work, we quantify this vulnerability, demonstrating that state-of-the-art VLAs such as $π_{0.5}$ and SmolVLA, suffer catastrophic performance degradation, dropping from 90\% success rates to as low as 2\%, under common signal artifacts. To mitigate this, we introduce the Corruption Restoration Transformer (CRT), a plug-and-play and model-agnostic vision transformer designed to immunize VLA models against sensor disturbances. Leveraging an adversarial training objective, CRT restores clean observations from corrupted inputs without requiring computationally expensive fine-tuning of the underlying model. Extensive experiments across the LIBERO and Meta-World benchmarks demonstrate that CRT effectively recovers lost performance, enabling VLAs to maintain near-baseline success rates, even under severe visual corruption.

机器人视觉修复鲁棒性VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。