arXiv:2409.03911cs.CVcs.AI2024-09ECCV

用生成模型提升加泰罗尼亚历史照片的自动描述能力

The Role of Generative Systems in Historical Photography Management: A Case Study on Catalan Archives

  • 基于视觉适应与语言接近性,优化跨语言图像描述模型
  • 在加泰罗尼亚档案历史照片上实现显著描述准确率提升
  • 为小语种历史图像管理提供可复用的迁移学习方案

遗产机构中图像分析用于自动化摄影管理的趋势日益增长。此类工具可降低手动标注新数据源的人力成本,同时通过在线索引和搜索引擎实现公民快速访问。然而,现有标记与描述工具多针对英文现代照片设计,忽视了少数语言的历史语料库,而这些语料库具有内在特殊性。本研究旨在量化生成系统在历史资料描述中的贡献,以加泰罗尼亚档案馆的历史照片字幕任务为案例。研究结果为从业者提供了基于视觉适应与语言相近性的图像描述模型迁移学习工具与方向。

原文摘要 · Abstract (English)

The use of image analysis in automated photography management is an increasing trend in heritage institutions. Such tools alleviate the human cost associated with the manual and expensive annotation of new data sources while facilitating fast access to the citizenship through online indexes and search engines. However, available tagging and description tools are usually designed around modern photographs in English, neglecting historical corpora in minoritized languages, each of which exhibits intrinsic particularities. The primary objective of this research is to study the quantitative contribution of generative systems in the description of historical sources. This is done by contextualizing the task of captioning historical photographs from the Catalan archives as a case study. Our findings provide practitioners with tools and directions on transfer learning for captioning models based on visual adaptation and linguistic proximity.

历史图像生成模型跨语言迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。