arXiv:2512.19675econ.GNcs.CV2025-12被引 3

用多模态大模型从古籍扫描图自动构建德国专利数据集

Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)

  • 用Gemini模型分析9562张历史扫描图,自动提取30万条专利信息
  • 生成数据质量优于人工,速度超人工795倍,成本降为1/205
  • 开源工具链,让非技术研究者也能快速复用

我们利用多模态大语言模型(LLMs)构建了1877至1918年间306,070项德国专利的数据集,基于9,562份档案图像扫描件,通过由Gemini-2.5-Pro和Gemini-2.5-Flash-Lite驱动的自动化流水线完成。基准测试显示,该方法生成的数据质量优于研究助理,同时在构建效率上超过人工795倍,成本降低至1/205。每页图像约含20至50条专利信息,以双栏排版,使用哥特体与罗马字体印刷,布局与字体复杂度表明传统方法难以应对。本研究证明多模态大模型是经济史数据构建范式变革的关键。我们已开源基准测试结果、专利数据集及基于LLM的处理流水线,可借助LLM辅助编码工具轻松适配其他图像语料库,降低非技术研究者的门槛。最后,我们分析了部署大模型进行历史数据构建的经济学问题,并探讨其对经济史领域的潜在影响。

原文摘要 · Abstract (English)

We leverage multimodal large language models (LLMs) to construct a dataset of 306,070 German patents (1877-1918) from 9,562 archival image scans using our LLM-based pipeline powered by Gemini-2.5-Pro and Gemini-2.5-Flash-Lite. Our benchmarking exercise provides tentative evidence that multimodal LLMs can create higher quality datasets than our research assistants, while also being more than 795 times faster and 205 times cheaper in constructing the patent dataset from our image corpus. About 20 to 50 patent entries are embedded on each page, arranged in a double-column format and printed in Gothic and Roman fonts. The font and layout complexity of our primary source material suggests to us that multimodal LLMs are a paradigm shift in how datasets are constructed in economic history. We open-source our benchmarking and patent datasets as well as our LLM-based data pipeline, which can be easily adapted to other image corpora using LLM-assisted coding tools, lowering the barriers for less technical researchers. Finally, we explain the economics of deploying LLMs for historical dataset construction and conclude by speculating on the potential implications for the field of economic history.

多模态大模型历史数据经济史文档解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。