arXiv:2606.10862cs.CVcs.AI2026-06被引 1

提出新基准与视角想象方法,提升机器人在遮挡下的操作鲁棒性。

LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination

论文配图:LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination
图 1 · 摘自论文原文
  • 通过生成补全视角,融合真实与想象观测进行决策。
  • 现有顶尖模型在遮挡下性能下降超40%。
  • 无需额外摄像头,适合真实场景机器人应用。

视觉-语言-动作(VLA)模型在标准操作基准上表现优异,但多数评估假设任务相关物体完全可见,这在现实场景中常不成立,因遮挡导致操作信息部分不可见。本文将场景诱导遮挡视为VLA模型的核心挑战,提出面向遮挡的LIBERO-Occ基准。实验表明,当前先进VLA模型在遮挡条件下性能显著下降。为此,我们提出视角想象(Viewpoint Imagination, VIM)机制:从主观测生成互补视角,并联合真实与想象证据进行动作预测。VIM在多个任务套件、遮挡类型及严重程度下均提升鲁棒性,且部署时无需额外摄像头,表明视角想象是部分可观测操作中感知补全的有效机制。相关基准与代码已开源。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-relevant objects are fully visible. This assumption often fails in realistic settings, where occlusion makes manipulation partially observable. In this paper, we study \textit{scene-induced occlusion} as a fundamental challenge for VLA models and introduce \textbf{LIBERO-Occ}, an occlusion-oriented extension of LIBERO. Experiments show that state-of-the-art VLAs suffer substantial performance degradation under occlusion. To address this issue, we propose \textbf{Viewpoint Imagination (VIM)}, which generates a complementary view from an occluded primary observation and conditions action prediction on both observed and imagined evidence. VIM improves robustness across task suites, occlusion types, and severity levels without requiring additional cameras at deployment time, suggesting that viewpoint imagination is an promising mechanism for perception completion in partially observable manipulation. Our benchmark and corresponding code are available at: \href{https://github.com/litsh/Libero-Occ}{https://github.com/litsh/Libero-Occ}.

机器人遮挡视觉语言动作感知补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。