arXiv:2608.05911cs.CV2026-08

用视觉语言模型重建20世纪亚美尼亚侨民商业广告地图

Mapping Armenian Paris: Extracting and Geocoding Commercial Advertisements from the 20th-Century Diaspora Press

  • 用VLM模型自动识别并解析曲面扫描的西亚美尼亚文广告
  • 构建500页含3270条广告的语料库,实现高精度地理标注
  • 方法可复用于其他资源匮乏的历史语言数据集

本文提出一个基于IIIF的端到端流程,将法国亚美尼亚20世纪报刊数字化为巴黎亚美尼亚商业社区的互动地图。通过视觉语言模型(VLMs)在每页中定位、识别并结构化商业广告,克服了传统行级CRNN OCR在强烈弯曲扫描图上失效的问题。该方法在未充分支持的西亚美尼亚文上实现高效数据获取,远超人工标注规模且保持可靠性。成果包括一个500页的西亚美尼亚文报刊语料库,包含3,270条广告级标注,一套可在Label Studio中单次标注完成检测与语义字段的模板,以及可复现的工作流,适用于其他资源匮乏的历史文本集合。研究证明,基于VLM的数据自举策略对(西)亚美尼亚语等历史语言具有显著效能。

原文摘要 · Abstract (English)

This paper presents an end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Parisian Armenian commercial community. On each page, commercial advertisements are located, read, and parsed into structured records, which are then geocoded and placed on the map. Western Armenian is under-resourced and unsupported by off-the-shelf layout and OCR models, so the pipeline uses vision-language models (VLMs) as a data-bootstrapping strategy: they produce usable structured records at a scale hand annotation could not reach, and stay reliable on the strongly curved scans where conventional line-level CRNN OCR breaks down. The contribution includes a 500-page Western Armenian press corpus with 3,270 advertisement-level annotations, a Label Studio template that captures detection and semantic fields in a single annotation pass, and a reproducible workflow transposable to other under-resourced historical corpora. More broadly, the work shows that VLM-driven data bootstrapping is an effective lever for under-resourced historical languages such as (Western) Armenian.

历史语言视觉语言模型地理标注数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。