arXiv:2605.16338cs.DLcs.CL2026-05

用AI自动化古籍数字化,让无标签数据变可检索。

Vidya: An AI-Driven Modular Pipeline for Archival Automation and Semantic Metadata Enrichment

论文配图:Vidya: An AI-Driven Modular Pipeline for Archival Automation and Semantic Metadata Enrichment
图 1 · 摘自论文原文
  • 用大模型+开源工具构建模块化流水线,自动补全档案元数据。
  • 处理速度从数十年缩短至数天,符合国际档案标准。
  • 适合图书馆、博物馆等机构低成本部署,支持开放共享。

历史档案的大规模数字化带来了悖论:大量数字资源因缺乏元数据而成为‘暗数据’。人工著录耗时费力,严重制约发现与复用。我们提出Vidya,一种由大语言模型(LLMs)与开源工具协同的模块化流水线,可规模化实现语义增强与档案入库。Vidya通过YAML定义的本体和Pydantic验证约束生成过程,将概率性输出转化为确定性的结构化JSON。该系统由庞塔格拉萨州立大学(UEPG)数字人文与创新实验室(LAMUHDI)开发,遵循创客理念与开源实践,仅需普通硬件即可在记忆机构低成本部署。我们对比了不同LLM表现,并进行成本效益分析,证明其能将处理时间从数十年压缩至数日,同时符合NOBRADE与ISAD(G)标准。

原文摘要 · Abstract (English)

The large-scale digitization of historical archives has created a paradox: "dark data"-digital objects lacking metadata for retrieval. Manual archival description is slow and expensive, limiting discovery and reuse. We propose Vidya, a modular pipeline that orchestrates Large Language Models (LLMs) and FOSS tools to automate semantic enrichment and archival ingestion at scale. Vidya constrains generations using YAML-defined ontologies and Pydantic validation, producing deterministic, structured JSON outputs from probabilistic models. Developed at Laboratory for Digital Humanities and Innovation (LAMUHDI) of the State University of Ponta Grossa (UEPG), Vidya applies Maker principles and open-source practices to enable low-cost deployment in memory institutions using modest hardware. We compare LLM performance and present a cost-benefit analysis showing major gains, reducing processing time from decades to days while complying with NOBRADE and ISAD(G).

档案自动化大模型应用开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。