arXiv:2605.11727cs.AIcs.CL2026-05

让视觉语言模型直接用原始传感器数据,提升复杂场景下的理解准确率。

Allegory of the Cave: Measurement-Grounded Vision-Language Learning

论文配图:Allegory of the Cave: Measurement-Grounded Vision-Language Learning
图 1 · 摘自论文原文
  • 用原始传感器测量数据替代RGB图像作为输入,更贴近真实物理信号。
  • 在低光、高动态范围等困难场景下,指标比基线提升4.46个百分点。
  • 适合研究多模态推理、感知建模或需要高鲁棒性的应用开发者。

视觉语言模型通常基于后处理的RGB图像进行推理,但RGB渲染过程可能丢失、压制或量化传感器原始证据。本文研究将视觉接口向底层相机测量信号靠近是否能提升模型的语义对齐能力。提出测量接地式视觉语言学习框架PRISM-VL,结合RAW-derived Meas.-XYZ输入、相机条件化对齐机制以及曝光分段监督聚合策略,实现从RGB代理任务向测量域观测的监督迁移。基于15万条质量可控的指令微调数据集和一个针对低光照、高动态范围、可见性敏感及幻觉敏感场景的独立测试集,PRISM-VL-8B在0.6120 BLEU、0.4571 ROUGE-L和82.66% LLM-Judge准确率上优于基线Qwen3-VL-8B,分别提升+0.1074 BLEU、+0.1071 ROUGE-L和+4.46个百分点。结果表明,部分视觉语言模型的对齐误差源于RGB渲染过程中的信息损失,保留测量域证据可有效增强多模态推理能力。

原文摘要 · Abstract (English)

Vision-language models typically reason over post-ISP RGB images, although RGB rendering can clip, suppress, or quantize sensor evidence before inference. We study whether grounding improves when the visual interface is moved closer to the underlying camera measurement. We formulate measurement-grounded vision-language learning and instantiate it as PRISM-VL, which combines RAW-derived Meas.-XYZ inputs, camera-conditioned grounding, and Exposure-Bracketed Supervision Aggregation for transferring supervision from RGB proxies to measurement-domain observations. Using a quality-controlled 150K instruction-tuning set and a held-out benchmark targeting low-light, HDR, visibility-sensitive, and hallucination-sensitive cases, PRISM-VL-8B reaches 0.6120 BLEU, 0.4571 ROUGE-L, and 82.66\% LLM-Judge accuracy, improving over the RGB Qwen3-VL-8B baseline by +0.1074 BLEU, +0.1071 ROUGE-L, and +4.46 percentage points. These results suggest that part of VLM grounding error arises from information lost during RGB rendering, and that preserving measurement-domain evidence can improve multimodal reasoning.

视觉语言模型原始数据多模态推理低光成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。