arXiv:2601.09433cs.CVcs.AI2026-01

用视觉Transformer识别古罗马钱币纹样,效果优于传统卷积网络。

Do Transformers Understand Ancient Roman Coin Motifs Better than CNNs?

  • 首次将ViT用于钱币语义元素自动识别,结合图像与文本数据端到端训练。
  • ViT在准确率上显著高于新训练的CNN模型,验证了其对复杂纹样的捕捉能力。
  • 适合历史考古、数字人文及多模态学习研究者参考。

古钱币的自动化分析有望帮助研究人员从大量钱币中提取更多历史信息,并协助收藏者理解所购藏品。近期研究已展示出利用卷积神经网络(CNN)识别钱币常见语义元素的潜力。本文首次将近期提出的视觉变换器(ViT)深度学习架构应用于钱币语义元素识别任务,采用图像与非结构化文本的多模态数据进行全自动学习。文章总结了该领域的前期研究,讨论了ViT与CNN模型在古钱币分析中的训练与实现,并评估了二者性能。结果表明,ViT模型在准确率上优于新训练的CNN模型。

原文摘要 · Abstract (English)

Automated analysis of ancient coins has the potential to help researchers extract more historical insights from large collections of coins and to help collectors understand what they are buying or selling. Recent research in this area has shown promise in focusing on identification of semantic elements as they are commonly depicted on ancient coins, by using convolutional neural networks (CNNs). This paper is the first to apply the recently proposed Vision Transformer (ViT) deep learning architecture to the task of identification of semantic elements on coins, using fully automatic learning from multi-modal data (images and unstructured text). This article summarises previous research in the area, discusses the training and implementation of ViT and CNN models for ancient coins analysis and provides an evaluation of their performance. The ViT models were found to outperform the newly trained CNN models in accuracy.

视觉Transformer古钱币分析多模态学习图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。