测试视觉语言模型读取仪表盘的能力,发现其表现远低于人类。
Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- 构建可生成可控仪表图像的合成管道,支持多种细节变化
- 主流VLM在真实与合成数据上均表现不佳,最高准确率不足60%
- 用强化学习微调显著提升性能,适合研究精确空间感知的学者
读取测量仪器对人类而言轻而易举,且无需太多领域知识,但当前视觉语言模型(VLMs)在此任务上仍面临挑战。本文提出MeasureBench,一个涵盖真实与合成图像的视觉测量阅读基准,包含可扩展的数据合成流程。该流程能程序化生成特定类型的刻度表,可控调节指针、刻度、字体、光照和杂乱程度等细节,实现大规模变异。对主流专有及开源VLM进行评估显示,即使最强模型在测量阅读任务中整体表现依然有限。初步实验表明,在合成数据上使用强化微调(RFT)可显著提升模型在域内合成集和真实图像上的性能。分析揭示当前VLM在细粒度空间定位方面存在根本性局限。我们希望本资源与代码发布能推动未来在视觉地物数值理解与精准空间感知方面的进展,弥合识别数字与测量世界之间的差距。
原文摘要 · Abstract (English)
Reading measurement instruments is effortless for humans and requires relatively little domain expertise, yet it remains surprisingly challenging for current vision-language models (VLMs) as we find in preliminary evaluation. In this work, we introduce MeasureBench, a benchmark on visual measurement reading covering both real-world and synthesized images of various types of measurements, along with an extensible pipeline for data synthesis. Our pipeline procedurally generates a specified type of gauge with controllable visual appearance, enabling scalable variation in key details such as pointers, scales, fonts, lighting, and clutter. Evaluation on popular proprietary and open-weight VLMs shows that even the strongest frontier VLMs struggle with measurement reading in general. We have also conducted preliminary experiments with reinforcement finetuning (RFT) over synthetic data, and find a significant improvement on both in-domain synthetic subset and real-world images. Our analysis highlights a fundamental limitation of current VLMs in fine-grained spatial grounding. We hope this resource and our code releases can help future advances on visually grounded numeracy and precise spatial perception of VLMs, bridging the gap between recognizing numbers and measuring the world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。