arXiv:2603.10340cs.CVcs.AI2026-03被引 2

通过概念门控去噪提升机器人在杂乱环境中的操作精度

Overcoming Visual Clutter in Vision Language Action Models via Concept-Gated Visual Distillation

  • 用指令分离安全区与干扰区,双重精炼目标定位
  • 傅里叶补全生成无干扰图像,保留关键几何结构
  • 无需训练即可显著提升复杂场景成功率至77.5%

视觉-语言-动作(VLA)模型虽具备出色零样本泛化能力,但在杂乱环境中常因背景干扰导致特征稀释,产生‘精度-推理差距’。为此,我们提出概念门控视觉蒸馏(CGVD),一种无需训练、通用性强的推理框架,可稳定VLA策略。CGVD将指令解析为安全区与干扰区,通过交叉验证与空间消歧双重机制,显式惩罚误检并锁定真实操作目标。随后利用基于傅里叶的图像修复技术处理场景,生成抑制语义干扰但保留关键空间几何与视觉本体感的干净观测。在高度杂乱的操作任务中,该方法有效防止性能下降。在密集语义干扰环境下,相较于基线43.0%的成功率,本方法达到77.5%。通过严格属性约束,证明推理时视觉蒸馏是实现鲁棒机器人操作的关键前提。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models demonstrate impressive zero-shot generalization but frequently suffer from a "Precision-Reasoning Gap" in cluttered environments. This failure is driven by background-induced feature dilution, where high-frequency semantic noise corrupts the geometric grounding required for precise manipulation. To bridge this gap, we propose Concept-Gated Visual Distillation (CGVD), a training-free, model-agnostic inference framework that stabilizes VLA policies. CGVD operates by parsing instructions into safe and distractor sets, utilizing a two-layer target refinement process--combining cross-validation and spatial disambiguation--to explicitly penalize false positives and isolate genuine manipulation targets. We then process the scene via Fourier-based inpainting, generating a clean observation that actively suppresses semantic distractors while preserving critical spatial geometry and visual proprioception. Extensive evaluations in highly cluttered manipulation tasks demonstrate that CGVD prevents performance collapse. In environments with dense semantic distractors, our method significantly outperforms state-of-the-art baselines, achieving a 77.5% success rate compared to the baseline's 43.0%. By enforcing strict attribute adherence, CGVD establishes inference-time visual distillation as a critical prerequisite for robust robotic manipulation in the clutter.

视觉推理机器人操作去噪零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。