解决对话式仪表读数中模型依赖外观、状态不一致的问题
State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading

- 设计三层对齐框架,强制模型关注仪表状态的几何结构
- 在可控钟表与仪表数据集上准确率提升超20%且抗光照变化能力增强
- 适合需要高精度工业仪表识别或视觉-语言对齐的研究者
多模态大模型在通用任务上表现优异,但在基于对话的仪表读数任务中仍显脆弱。本文通过受控基准和特征空间探测发现,当前模型不仅在仪表读数上准确率不理想,即使在视角和光照变化下(本质状态不变),性能也急剧下降。探查分析表明,相同状态样本在外观变化下未被一致聚类,邻近状态也未能保持连续值所隐含的局部结构。这说明现有模型忽视了仪表状态的内在几何关系,而依赖表面外观线索。为此,我们提出三层次状态一致性对齐框架TriSCA,包含:状态距离感知表示对齐、元数据引导的观测到状态监督、以及状态感知的目标对齐。在受控时钟与仪表基准上进行大量消融实验和评估,并在外部真实世界基准上验证,结果证明该方法有效。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have achieved impressive progress on general multimodal tasks, yet they remain brittle on dial-based measurement reading. In this paper, we study this problem through controlled benchmarks and feature-space probing, and show that current MLLMs not only achieve unsatisfactory accuracy on dial-based readout, but also suffer sharp performance drops under viewpoint and illumination changes even when the underlying dial state remains fixed. Our probing analysis further reveals that same-state samples under appearance variation are not consistently clustered, while neighboring states fail to preserve the local structure implied by continuous dial values. These findings suggest that existing MLLMs largely ignore the intrinsic state geometry of dial measurement tasks and instead rely on superficial appearance cues. Motivated by this diagnosis, we propose TriSCA, a tri-level state-consistent alignment framework for dial-based measurement reading. Specifically, TriSCA consists of state-distance-aware representation alignment, metadata-grounded observation-to-state supervision, and state-aware objective alignment. Extensive ablation studies and evaluation experiments on controlled clock and gauge benchmarks, together with evaluation on an external real-world benchmark, demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。