arXiv:2603.08011cs.CV2026-03中稿 · CVPR

提升视觉语言模型读取真实场景时钟的能力,解决指针混淆与环境干扰问题。

It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models

  • 构建真实场景时钟数据集TickTockVQA,包含丰富背景与时间标注。
  • 通过Swap-DPO微调,使模型在遮挡、光照变化下准确识别时分针。
  • 适合关注视觉时空推理与多模态理解的研究者使用。

尽管视觉语言模型在复杂多模态推理任务中取得显著进展,但其在真实场景中读取模拟时钟仍面临巨大挑战。现有时钟数据集多为合成或平面化,缺乏样式多样性与真实背景,导致模型在训练后表现出薄弱的时空推理能力,常混淆时针与分针,且在遮挡、光照变化和杂乱背景等常见条件下表现不佳。为此,我们提出人工标注的TickTockVQA数据集,涵盖多样真实场景,提供明确的时/分标注,并在视觉上下文可推断时添加AM/PM标签。同时,我们设计基于直接偏好优化的Swap-DPO微调框架,引导模型更准确地进行时间解读。实验表明,该方法显著提升了模型在真实条件下的读时准确性与鲁棒性,为未来视觉语言模型的时空推理与视觉理解研究奠定基础。

原文摘要 · Abstract (English)

Advances in vision-language models (VLMs) have achieved remarkable success on complex multimodal reasoning tasks, leading to the assumption that they should also excel at reading analog clocks. However, contrary to this expectation, our study reveals that reading analog clocks in real-world environments remains a significant challenge for state-of-the-art VLMs. Existing analog clock datasets are largely synthetic or planar with limited stylistic diversity and minimal background context, failing to capture the visual variability of real-world scenes. As a result, VLMs trained on such data exhibit weak spatiotemporal reasoning, frequently confusing the hour and minute hands and struggling under common visual conditions such as occlusion, lighting variation, and cluttered backgrounds. To address this issue, we introduce TickTockVQA, a human-annotated dataset containing analog clocks in diverse real-world scenarios. TickTockVQA provides explicit hour and minute annotations, and includes an AM/PM tag when it is inferable from the visual context. Furthermore, we propose Swap-DPO, a direct preference optimization-based fine-tuning framework to align model reasoning toward accurate time interpretation. Experimental results demonstrate that our approach substantially enhances clock reading accuracy and robustness under real-world conditions, establishing a foundation for future research on spatiotemporal reasoning and visual understanding in VLMs.

时钟识别视觉语言模型时空推理数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。