arXiv:2512.24985cs.CVcs.AI2025-12

评测视觉语言模型在暗光室内场景下的感知能力,揭示其在低照度下的性能退化。

DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes

  • 构建物理真实感的暗光图像合成流程,模拟真实光照与传感器噪声。
  • 包含9400对可验证的问答对,覆盖五类视觉基本元素,支持多级暗光测试。
  • 发现主流VLM在暗光下普遍表现下降,低光增强方法效果不稳定。

视觉语言模型(VLMs)正被广泛用作具身智能体的核心推理模块。现有基准主要评估模型在理想光照条件下的表现,但全天候运行需应对包括夜间或暗环境在内的多种视觉退化问题,这一核心挑战长期被忽视。为此,我们提出DarkQA,一个开源基准,用于评估具身场景中多层级暗光条件下的感知基础能力。DarkQA通过单视角第一人称观察,在可控退化水平下评估视觉感知失败,避免其与复杂任务耦合。该基准包含9400对确定性生成且可验证的问答对,涵盖五类视觉基本元素。其关键设计是物理保真度:退化过程在原始线性空间建模,模拟基于物理的光照衰减与传感器噪声,并经类ISP渲染流水线处理;我们进一步用真实配对低光数据验证了合成效果。我们评估了代表性VLMs和低光图像增强(LLIE)预处理方法。结果表明,所有VLM在弱光与噪声下均出现持续性能下降,而LLIE仅提供依赖严重程度但不稳定的恢复效果。通过系统评估多种前沿VLM与LLIE模型,我们揭示了它们在严苛视觉条件下的局限性。代码与数据集将在论文录用后公开。项目网站:https://darkqa-benchmark.github.io

原文摘要 · Abstract (English)

Vision Language Models (VLMs) are increasingly adopted as central reasoning modules for embodied agents. Existing benchmarks evaluate their capabilities under ideal, well-lit conditions, yet robust 24/7 operation demands performance under a wide range of visual degradations, including low-light conditions at night or in dark environments, a core necessity that has been largely overlooked. To address this underexplored challenge, we present DarkQA, an open-source benchmark for evaluating perceptual primitives under multi-level low-light conditions in embodied scenarios. DarkQA evaluates single-view egocentric observations across controlled degradation levels, isolating low-light perceptual failures before they are entangled with complex embodied tasks. The benchmark contains 9.4K deterministically generated and verifiable question-image pairs spanning five visual-primitive families. A key design feature of DarkQA is its physical fidelity: visual degradations are modeled in linear RAW space, simulating physics-based illumination drop and sensor noise followed by an ISP-inspired rendering pipeline; we further validate the synthesis against real paired low-light camera data. We evaluate representative VLMs and Low-Light Image Enhancement (LLIE) preprocessing methods. Results show consistent VLM degradation under low illumination and sensor noise, while LLIE provides severity-dependent but unstable recovery. We demonstrate the utility of DarkQA by evaluating a wide range of state-of-the-art VLMs and Low-Light Image Enhancement (LLIE) models, and systematically reveal VLMs' limitations when operating under these challenging visual conditions. Our code and benchmark dataset will be released upon acceptance. Project website: https://darkqa-benchmark.github.io

视觉语言模型暗光检测具身智能基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。