arXiv:2512.03992cs.CVcs.AI2025-12

提出新基准与框架,评估视觉语言模型在持续干扰下的可靠性与伦理一致性。

Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness

  • 用对抗性视觉扰动模拟真实环境,评估模型连续推理中的错误传播。
  • 在DIQ-H基准上,模型准确率从72.2%提升至83.3%,相对改善15.3%。
  • 适合关注机器人、自动驾驶等安全关键系统的开发者和研究者。

视觉语言模型(VLMs)在具身智能与安全关键应用中至关重要,如机器人与自动驾驶系统。然而现有基准多聚焦静态或精心筛选的视觉输入,忽视了对抗性条件、价值错位及持续部署中的误差累积问题。当前评估或忽略真实世界扰动,或未考虑推理不一致的长期影响。为此,我们提出首个针对连续序列中对抗性视觉条件的评估基准——退化图像质量导致幻觉(DIQ-H),模拟运动模糊、传感器噪声与压缩伪影等现实压力,衡量这些污染如何引发持续性错误与时间上的价值错位。该基准显式建模误差传播及其长期价值一致性。为提升评估可扩展性并降低安全评估成本,我们提出价值引导迭代精炼(VIR)框架,利用轻量级VLM自动检测并修正价值错位,将准确率从72.2%提升至83.3%,相对提高15.3%。DIQ-H与VIR共同构成具身智能安全评估的可靠平台,揭示了模型在错误恢复、伦理一致性与时间价值对齐方面的脆弱性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are essential for embodied AI and safety-critical applications, such as robotics and autonomous systems. However, existing benchmarks primarily focus on static or curated visual inputs, neglecting the challenges posed by adversarial conditions, value misalignment, and error propagation in continuous deployment. Current benchmarks either overlook the impact of real-world perturbations, or fail to account for the cumulative effect of inconsistent reasoning over time. To address these gaps, we introduce the Degraded Image Quality Leading to Hallucinations (DIQ-H) benchmark, the first to evaluate VLMs under adversarial visual conditions in continuous sequences. DIQ-H simulates real-world stressors including motion blur, sensor noise, and compression artifacts, and measures how these corruptions lead to persistent errors and misaligned outputs across time. The benchmark explicitly models error propagation and its long-term value consistency. To enhance scalability and reduce costs for safety-critical evaluation, we propose the Value-Guided Iterative Refinement (VIR) framework, which automates the generation of high-quality, ethically aligned ground truth annotations. VGIR leverages lightweight VLMs to detect and refine value misalignment, improving accuracy from 72.2% to 83.3%, representing a 15.3% relative improvement. The DIQ-H benchmark and VGIR framework provide a robust platform for embodied AI safety assessment, revealing vulnerabilities in error recovery, ethical consistency, and temporal value alignment.

视觉语言模型安全评估误差传播伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。