arXiv:2607.10655cs.RO2026-07

让机器人模型聚焦关键视觉区域,减少误判,提升泛化能力。

Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models

论文配图:Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models
图 1 · 摘自论文原文
  • 通过生成任务相关的视觉掩码,引导模型关注真正重要的区域。
  • 减少微调时间,抑制过拟合,在环境变化下仍保持稳定表现。
  • 无需改动主模型结构,适配各类机器人基础模型使用。

机器人基础模型在多任务、跨平台和语言控制方面取得显著进展,但其在真实场景中的鲁棒部署仍面临挑战,部分原因在于策略常混淆因果相关与虚假的场景关联。我们将其归因于‘捷径学习’:过度依赖训练数据中可预测但非因果的关联,而非决定成功动作的任务相关视觉证据。针对此问题,提出人工中心视野感知(AFP),一种轻量级、与策略无关的模块,接收与视觉-语言-动作或世界动作模型相同的输入,生成任务条件下的关键区域掩码(包括物体、机器人及其他行动相关区域)。该掩码主要作为微调阶段的辅助对齐信号,引导策略注意力聚焦任务相关区域,不改变核心架构。微调后,策略在原始观测流上执行,无需AFP参与控制循环。在多个先进机器人基础模型上评估显示,引入中心视野感知可缩短微调时间,抑制过拟合,并提升环境扰动下的泛化性能。掩码质量与对齐损失设计的消融实验进一步表明,性能提升源于引导学习聚焦任务相关视觉证据。结果表明,任务条件下的中心视野感知是提升机器人基础模型鲁棒性、数据效率与可扩展性的实用机制。

原文摘要 · Abstract (English)

Robotic foundation models have recently made substantial progress in multi-task capability, cross-embodiment transfer, and language-conditioned control. Yet robust deployment across diverse real-world settings remains difficult, in part because policies often fail to distinguish causally relevant visual structure from spurious scene-level correlations. We identify this failure mode as shortcut learning: the tendency to exploit predictive but non-causal correlations in the training distribution rather than the task-relevant visual evidence that determines successful action. Although shortcut learning has been extensively studied in computer vision and broader machine learning, its role in robotic foundation models remains comparatively underexplored. We propose Artificial Foveated Perception (AFP), a lightweight, policy-agnostic module that takes the same vision and language inputs as Vision-Language-Action and World Action Model pipelines and predicts task-conditioned masks over relevant objects, the robot, and other action-critical regions. We use these masks primarily as an auxiliary grounding signal during fine-tuning, aligning policy attention with task-relevant regions while leaving the core architecture unchanged. After fine-tuning, the policy executes on the original observation stream without requiring AFP in the control loop. We evaluate AFP across state-of-the-art robotic foundation models and show that foveated perception reduces fine-tuning time, suppresses overfitting, and improves generalization under environmental perturbations. Ablations over mask quality and grounding-loss design further show that these gains arise from directing policy learning toward task-relevant visual evidence. These results suggest that task-conditioned foveated perception is a practical mechanism for making robotic foundation models more robust, data-efficient, and scalable.

机器人视觉注意泛化能力模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。