用多模型协作自动生成高质量档案元数据
Automated Archival Descriptions with Federated Intelligence of LLMs
- 通过联邦优化整合多个大模型智能生成档案元数据
- 相比单模型方案,元数据质量和可靠性显著提升
- 适合需要标准化档案处理的图书馆与档案馆
制定档案标准需要专业知识,而手动为档案材料创建元数据既耗时又易出错。本文探索代理型AI与大语言模型(LLMs)在实现标准化档案描述流程中的潜力。提出一种基于代理AI的自动化系统,可生成高质量档案材料的元数据。设计了一种联邦优化方法,融合多个LLMs的智能以构建最优元数据。同时提出应对大模型生成一致性挑战的方法。在涵盖多种文档类型和格式的真实世界档案数据集上进行了广泛实验。评估结果验证了该技术的可行性,并表明联邦优化方法在元数据质量与可靠性方面优于单模型解决方案。
原文摘要 · Abstract (English)
Enforcing archival standards requires specialized expertise, and manually creating metadata descriptions for archival materials is a tedious and error-prone task. This work aims at exploring the potential of agentic AI and large language models (LLMs) in addressing the challenges of implementing a standardized archival description process. To this end, we introduce an agentic AI-driven system for automated generation of high-quality metadata descriptions of archival materials. We develop a federated optimization approach that unites the intelligence of multiple LLMs to construct optimal archival metadata. We also suggest methods to overcome the challenges associated with using LLMs for consistent metadata generation. To evaluate the feasibility and effectiveness of our techniques, we conducted extensive experiments using a real-world dataset of archival materials, which covers a variety of document types and formats. The evaluation results demonstrate the feasibility of our techniques and highlight the superior performance of the federated optimization approach compared to single-model solutions in metadata quality and reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。