arXiv:2502.05092cs.CVcs.AI2025-02中稿 · ICLR被引 13

测试大模型看钟表和日历的能力,发现仍难准确理解时间。

Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs

  • 构建钟表和日历问答数据集,覆盖多种样式与复杂问题。
  • 模型在识别时钟指针和计算日期时错误率高,尤其涉及推理时。
  • 适合研究视觉时序理解或评估多模态模型认知能力的学者。

理解视觉中的时间是基本认知技能,但对多模态大语言模型(MLLMs)仍是挑战。本文通过分析模拟时钟和年度日历图像,评估MLLMs在视觉识别、数值推理和时序推断方面的能力。为此,我们构建了两个子集:1)ClockQA,包含标准、黑盘、无秒针、罗马数字、箭头指针等风格时钟,搭配时间相关问题;2)CalendarQA,由年度日历图像组成,问题涵盖已知日期(如圣诞节、元旦)及需计算的日期(如一年中的第100天或第153天)。实验表明,尽管近期有进展,可靠理解时间对MLLMs而言仍是重大挑战。

原文摘要 · Abstract (English)

Understanding time from visual representations is a fundamental cognitive skill, yet it remains a challenge for multimodal large language models (MLLMs). In this work, we investigate the capabilities of MLLMs in interpreting time and date through analogue clocks and yearly calendars. To facilitate this, we curated a structured dataset comprising two subsets: 1) $\textit{ClockQA}$, which comprises various types of clock styles$-$standard, black-dial, no-second-hand, Roman numeral, and arrow-hand clocks$-$paired with time related questions; and 2) $\textit{CalendarQA}$, which consists of yearly calendar images with questions ranging from commonly known dates (e.g., Christmas, New Year's Day) to computationally derived ones (e.g., the 100th or 153rd day of the year). We aim to analyse how MLLMs can perform visual recognition, numerical reasoning, and temporal inference when presented with time-related visual data. Our evaluations show that despite recent advancements, reliably understanding time remains a significant challenge for MLLMs.

多模态时间理解视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。