arXiv:2604.22829cs.CV2026-04被引 1

测试发现大模型读动态仪表时表现不佳,难以可靠识别指针轨迹。

Lost in Motion: Vision Language Models Fail the Dynamic Gauges Test

  • 构建三类仪表的动态视频数据集,模拟不同运动速度
  • 主流大模型在指针轨迹理解上准确率不足,无法满足安全标准
  • 适用于工业自动化中对仪表读数可靠性要求高的场景

工业制造的数字化转型依赖自主机器人与传统模拟仪表交互的能力。视觉语言模型(VLMs)虽在仪表识别上表现良好,但在实时、精确的仪表读数任务中仍面临挑战。本文评估了GPT-5(5.4 Thinking和5.3 Instant)及Gemini 3(Pro和Flash)系列前沿模型,在包含圆形、线性与游标三种仪表的动态视频测试集上的表现。该数据集涵盖多种运动与速度模式。结果显示,当前顶级VLMs在理解指针轨迹和刻度语义方面能力有限,难以提供安全监控所需的可追溯性与可靠性。研究指出,这些模型尚未达到现有IEEE与ISO标准下被认定为可信合成仪表的性能门槛。

原文摘要 · Abstract (English)

The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, particularly analog gauges. Vision-Language Models (VLMs) have the potential to provide a general solution for gauge reading and have already shown good performance in instrument recognition. However, performing accurate, real-time gauge readings is a more complex task. This paper evaluates state-of-the-art models, including versions from the GPT-5 (5.4 Thinking and 5.3 Instant) and Gemini 3 (Pro and Flash) families, against a set of simple but realistic dynamic gauge reading scenarios. To facilitate this evaluation, we introduce a novel dataset comprising video sequences of three instruments of different gauge types: circular, linear, and Vernier, under diverse motion and speed profiles. Our findings indicate that the evaluated frontier VLMs, under our specific testing conditions, exhibit a limited ability to interpret needle trajectories and scale semantics, failing to provide the traceability and reliability needed for safety-critical monitoring. The results demonstrate that these models have not yet achieved the performance necessary to be classified as trustworthy synthetic instruments under existing IEEE and ISO standards.

视觉语言模型工业自动化仪表识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。