用大模型为长文本文献补充语义元数据,提升数字馆藏可检索性。
Metadata Enrichment of Long Text Documents using Large Language Models
- 结合人工与大模型,对1920-2020年长文本进行语义元数据增强
- 显著改善元数据缺失严重的数字馆藏的搜索效果与可访问性
- 适合数字人文、信息科学等领域的研究者使用
本项目通过人工与大语言模型相结合的方式,对来自HathiTrust数字图书馆、1920至2020年间发表的英文长文本文档(如论文和学位论文)进行了语义层面的元数据丰富与增强。该数据集为计算社会科学、数字人文和信息科学等领域提供了重要资源。研究表明,利用大模型增强元数据,可为数字馆藏引入原本未预设的多维度检索入口,尤其适用于现有元数据存在显著缺失的情况,有效提升检索准确率与资源可及性。
原文摘要 · Abstract (English)
In this project, we semantically enriched and enhanced the metadata of long text documents, theses and dissertations, retrieved from the HathiTrust Digital Library in English published from 1920 to 2020 through a combination of manual efforts and large language models. This dataset provides a valuable resource for advancing research in areas such as computational social science, digital humanities, and information science. Our paper shows that enriching metadata using LLMs is particularly beneficial for digital repositories by introducing additional metadata access points that may not have originally been foreseen to accommodate various content types. This approach is particularly effective for repositories that have significant missing data in their existing metadata fields, enhancing search results and improving the accessibility of the digital repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。