arXiv:2608.17723cs.CV2026-08

用视觉语言模型直接读取模拟仪表,无需复杂标注流程。

Vision-Language Models for Analog Gauge Reading: An Empirical Study of Specialization, Transfer and Reliability

论文配图:Vision-Language Models for Analog Gauge Reading: An Empirical Study of Specialization, Transfer and Reliability
图 1 · 摘自论文原文
  • 用QLoRA微调通用视觉语言模型,实现零样本到小样本的直接读数。
  • 合成数据上误差低至2.39%,工业数据上为4.43%。
  • 模型对模糊敏感,高置信度仍可能出错,适合安全关键场景需验证。

模拟仪表在工业环境中仍广泛存在,人工读数成本高且危险。本文研究直接从单目标仪表图像中读取数值的任务,贡献在于系统评估通用视觉语言模型(VLM)在专业化、迁移性、鲁棒性和可靠性方面的表现,无需显式的指针分割或几何读数流程。采用Qwen2.5-VL-7B-Instruct模型,在公开合成数据集、视频生成的Pressure Gauge数据集及专有工业数据集上,通过零样本提示、上下文学习(ICL)和量化低秩适应(QLoRA)进行参数高效微调。所有微调实验均采用固定20轮协议,最后一轮用于分析;对比了有无提供量程的模型以排除提示干扰。主要指标为归一化均百分比误差(MPE)。最佳微调结果:合成数据集上为2.39%(95%置信区间1.43-3.90%),Pressure Gauge数据集上为2.61%(CI 1.66-3.80%),专有工业数据集上为4.43%(CI 2.31-7.14%)。留一数据集实验显示显著迁移性能下降,鲁棒性测试表明高斯模糊是最强破坏因素。可靠性分析发现高置信度下仍可能发生错误,推动在安全关键场景中引入拒绝回答与独立验证机制。结果支持基于QLoRA的专用VLM用于单仪表读数,但尚不足以部署为完整工厂监控系统。

原文摘要 · Abstract (English)

Analog gauges remain common in industrial environments where manual inspection is costly or hazardous. The engineering application addressed here is direct numerical reading of single-target analog-gauge images, while the artificial-intelligence contribution is a systematic evaluation of specialization, transfer, robustness and reliability for a general-purpose vision-language model (VLM) without an explicit pointer-segmentation and geometric-reading pipeline. The Qwen2.5-VL-7B-Instruct model is evaluated using zero-shot prompting, in-context learning (ICL) and parameter-efficient fine-tuning with Quantized Low-Rank Adaptation (QLoRA) on a public synthetic dataset, a video-derived Pressure Gauge dataset and a proprietary industrial dataset. All fine-tuning experiments use a fixed 20-epoch protocol with the final epoch used for analysis; separate models with and without supplied gauge ranges remove prompt-setting confounds. The primary metric is range-normalized mean percentage error (MPE). The best fine-tuned MPE values are 2.39% on the synthetic dataset, with a 95% bootstrap confidence interval (CI) of 1.43-3.90%; 2.61% on the Pressure Gauge dataset, with a CI of 1.66-3.80%; and 4.43% on the proprietary industrial dataset, with a CI of 2.31-7.14%. Leave-one-dataset-out experiments reveal substantial transfer degradation on held-out synthetic and proprietary data, while robustness tests identify Gaussian blur as the strongest tested corruption. Reliability analysis shows that high-confidence errors remain possible, motivating abstention and independent validation in safety-critical use. These results support QLoRA-specialized VLMs for direct single-gauge reading but not yet a deployment-ready plant-monitoring pipeline.

视觉语言模型仪表识别工业检测可靠性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。