arXiv:2505.10862cs.CL2025-05被引 3

测试GPT-4.1读模拟钟表时间能力,发现其依赖训练数据模式而非真正理解。

Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?

  • 用不同钟面测试模型,检验其是否具备泛化能力
  • 模型在特定钟表上表现良好,但换类型就出错
  • 适合研究多模态模型泛化局限的读者

能够回答图像复杂问题的多模态大语言模型在读取模拟钟表时间时表现不佳,可能是因为训练数据中缺乏不同时刻的钟表图像。本文以最新MMLM GPT-4.1为对象,探究其失败原因及微调是否能解决该问题。结果表明,模型在读取时间方面有所进步,但其表现可能源于对训练数据中模式的依赖,而非真正理解时间读取机制。通过引入不同类型的钟表进行测试,揭示了模型在抽象与泛化方面的局限性。

原文摘要 · Abstract (English)

Multimodal Large Language Models which can answer complex questions on an image struggle to tell the time on analog clocks. This is probably due to the lack of images with clocks at different times in their training set. In this work we explore this issue with one of the latest MLLMs: GPT-4.1 to understand why MLLMs fail to tell the time and whether fine-tuning can solve the problem. The results show how models are making progress in reading the time on analog clocks. But have they really learned to do it, or have they only learned patterns in their training datasets? In this work we put the models to the test with different clocks to illustrate the limitations of MLLMs to abstract and generalize.

多模态模型时钟识别泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。