arXiv:2606.00828cs.CV2026-06被引 1

为机器人视觉语言模型设计真实物理环境下的抗干扰测试基准

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes

论文配图:RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
图 1 · 摘自论文原文
  • 从逆图形学出发,将视觉压力分解为材质、视角、光照、几何四维度
  • 实测发现不同物理因素会分别削弱视觉识别、推理与规划能力
  • 提出感知压力的智能代理,可主动修复图像后推理,提升复杂场景鲁棒性

视觉语言模型(VLM)在具身AI系统中展现出强大的视觉理解能力,但现有评测多基于干净图像或孤立扰动,未能覆盖真实物理场景形成的综合视觉压力。这导致评测范围狭窄且部分扰动在实际中罕见。为此,本文从逆图形学视角出发,提出RoboStressBench——一个评估具身场景中VLM对物理视觉压力鲁棒性的基准。受物理渲染方程启发,该基准将视觉压力分解为材料(M)、视角(V)、光照(L)和几何(G)四个物理基础维度,覆盖真实环境中的多样化视觉应力,同时支持对视觉识别、推理与规划等能力影响的可控分析。对主流VLM的全面评估揭示了压力特异性的失败模式,并发现不同物理因素对具身能力的损害各异,常被整体准确率掩盖。进一步提出一种压力感知的智能体求解器,可在推理前检测视觉压力并调用图像编辑技能,显著提升高压力场景下的表现。RoboStressBench为诊断与改进真实物理压力下的VLM感知提供了原则性框架,助力更可靠的具身AI系统发展。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception under real conditions is essential. However, existing benchmarks assess VLMs using clean images or isolated perturbations rather than stresses caused by physical scene formation. This design has two limitations: it covers only a narrow subset of everyday visual stresses, and some perturbations rarely appear in realistic embodied scenes. This gap raises a fundamental question: how can we define visual stress in a principled way that captures the diverse factors encountered in physical environments? To address this question, we formulate visual perception from an inverse graphics perspective and introduce RoboStressBench, a benchmark for evaluating VLM robustness to physical visual stress in embodied scenes. Inspired by the physical rendering equation, RoboStressBench decomposes visual stress into four physically grounded dimensions: Material (M), Viewpoint (V), Lighting (L), and Geometry (G). This design enables RoboStressBench to cover a broad range of visual stresses in real-world environments, while allowing controlled analysis of their effects on VLM capabilities such as visual recognition, reasoning, and planning. Through comprehensive evaluations of state-of-the-art VLMs, we identify stress-specific failure modes and reveal that different physical factors degrade different embodied capabilities, which are often obscured by aggregate accuracy. We further introduce a stress-aware agentic solver that detects visual stressors and invokes visual-editing skills before reasoning, improving robustness in high-stress scenarios. Overall, RoboStressBench provides a principled evaluation framework for diagnosing and improving VLM perception under real-world physical stress, supporting the development of more reliable embodied AI systems.

具身智能视觉压力模型评测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。