首个系统量化机器人视觉偏见的基准,揭示视角与色彩对决策影响
RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- 构建隔离变量的结构化任务生成框架,精准测量单因素及交互作用下的视觉偏见
- 三类智能体均存在显著偏见,视角是关键因素,高饱和色彩下成功率最高
- 视角会显著放大色彩偏见,引入语义对齐层可降低54.5%偏见
具身智能体的安全与可靠性依赖于准确无偏的视觉感知。然而现有基准多关注扰动下的泛化与鲁棒性,对视觉偏见的系统性量化仍显不足,限制了对感知如何影响决策稳定性的深入理解。为此,我们提出RoboView-Bias,首个专为系统量化机器人操作中视觉偏见而设计的基准,遵循因子隔离原则。通过结构化的变体生成框架与感知公平性验证协议,构建了2,127个任务实例,可稳健测量单一视觉因素及其交互作用引发的偏见。基于此基准,我们系统评估了三种代表性具身智能体在两种主流范式下的表现,发现:(i) 所有智能体均存在显著视觉偏见,其中相机视角是最关键因素;(ii) 智能体在高饱和度颜色下取得最高成功率,表明其继承自底层视觉语言模型(VLM)的视觉偏好;(iii) 视觉偏见呈现强非对称耦合,视角显著放大与色彩相关的偏见。最后,我们展示一种基于语义对齐层的缓解策略,在MOKA上使视觉偏见降低约54.5%。结果表明,系统分析视觉偏见是发展安全可靠通用具身智能体的前提。
原文摘要 · Abstract (English)
The safety and reliability of embodied agents rely on accurate and unbiased visual perception. However, existing benchmarks mainly emphasize generalization and robustness under perturbations, while systematic quantification of visual bias remains scarce. This gap limits a deeper understanding of how perception influences decision-making stability. To address this issue, we propose RoboView-Bias, the first benchmark specifically designed to systematically quantify visual bias in robotic manipulation, following a principle of factor isolation. Leveraging a structured variant-generation framework and a perceptual-fairness validation protocol, we create 2,127 task instances that enable robust measurement of biases induced by individual visual factors and their interactions. Using this benchmark, we systematically evaluate three representative embodied agents across two prevailing paradigms and report three key findings: (i) all agents exhibit significant visual biases, with camera viewpoint being the most critical factor; (ii) agents achieve their highest success rates on highly saturated colors, indicating inherited visual preferences from underlying VLMs; and (iii) visual biases show strong, asymmetric coupling, with viewpoint strongly amplifying color-related bias. Finally, we demonstrate that a mitigation strategy based on a semantic grounding layer substantially reduces visual bias by approximately 54.5\% on MOKA. Our results highlight that systematic analysis of visual bias is a prerequisite for developing safe and reliable general-purpose embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。