arXiv:2511.21982cs.CVcs.AI2025-11被引 2

用大模型提升指针仪表读数准确率,解决反光遮挡等难题

DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models

  • 通过注入几何与因果关系,让模型理解指针与刻度的物理关联
  • 在10730张真实图像上实现高精度读数,显著优于现有方法
  • 适合智能电网、工业检测等领域开发者使用

精确读取指针仪表数值在智能电力系统中至关重要,但现有方法因反光、遮挡、动态视角以及指针与刻度线过细等问题仍显脆弱。目前该领域缺乏大规模数据集支持鲁棒算法开发。为此,本文首次提出一个大规模基准数据集RPM-10K,包含10730张全面反映上述挑战的仪表图像。基于此数据集,我们提出一种新型视觉-语言模型MRLM,通过物理关系注入实现精准读数。MRLM不依赖图像级相关性学习,而是显式编码指针与刻度间的几何与因果关系,结合交叉注意力融合与自适应专家选择机制,使模型能理解仪表布局并生成精确数值。大量实验充分验证了该框架在新基准上的有效性。数据集与源码将公开于https://github.com/Event-AHU/DialBench。

原文摘要 · Abstract (English)

The precise reading recognition of pointer meters plays a key role in smart power systems, but existing approaches remain fragile due to challenges like reflections, occlusions, dynamic viewing angles, and overly between thin pointers and scale markings. Up to now, this area still lacks large-scale datasets to support the development of robust algorithms. To address these challenges, this paper first presents a new large-scale benchmark dataset for dial reading, termed RPM-10K, which contains 10730 meter images that fully reflect the aforementioned key challenges. Built upon the dataset, we propose a novel vision-language model for pointer meter reading recognition, termed MRLM, based on physical relation injection. Instead of exhaustively learning image-level correlations, MRLM explicitly encodes the geometric and causal relationships between the pointer and the scale, aligning perception with physical reasoning in the spirit of world-model perspectives. Through cross-attentional fusion and adaptive expert selection, the model learns to interpret dial configurations and generate precise numeric readings. Extensive experiments fully validated the effectiveness of our proposed framework on the newly proposed benchmark dataset. Both the dataset and source code will be released on https://github.com/Event-AHU/DialBench

仪表识别视觉语言模型物理推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。