用文学理解提升历史文献检索,让数字档案更包容
Cultural Analytics for Good: Building Inclusive Evaluation Frameworks for Historical IR
- 结合专家设计与大模型,构建可扩展的评估框架
- 利用19世纪小说语义提升非虚构文献检索准确率
- 适合关注数字人文与公平知识获取的研究者
本文融合信息检索与文化分析,推动历史知识的公平获取。基于英国图书馆BL19数字藏品(1700-1899年,超35,000件作品),我们构建了用于研究19世纪虚构与非虚构文本中语言、术语及检索变化的基准。方法结合专家驱动的查询设计、段落级相关性标注与大语言模型辅助,建立以人类经验为基础的可扩展评估框架。重点探索从虚构作品到非虚构作品的知识迁移,研究叙事理解与语义丰富性如何提升学术性事实材料的检索效果。该跨学科框架不仅提高检索准确性,还增强可解释性、透明度与文化包容性。本工作提供实用评估资源与方法论范式,助力开发支持历史意识的检索系统,最终构建更具解放性的知识基础设施。
原文摘要 · Abstract (English)
This work bridges the fields of information retrieval and cultural analytics to support equitable access to historical knowledge. Using the British Library BL19 digital collection (more than 35,000 works from 1700-1899), we construct a benchmark for studying changes in language, terminology and retrieval in the 19th-century fiction and non-fiction. Our approach combines expert-driven query design, paragraph-level relevance annotation, and Large Language Model (LLM) assistance to create a scalable evaluation framework grounded in human expertise. We focus on knowledge transfer from fiction to non-fiction, investigating how narrative understanding and semantic richness in fiction can improve retrieval for scholarly and factual materials. This interdisciplinary framework not only improves retrieval accuracy but also fosters interpretability, transparency, and cultural inclusivity in digital archives. Our work provides both practical evaluation resources and a methodological paradigm for developing retrieval systems that support richer, historically aware engagement with digital archives, ultimately working towards more emancipatory knowledge infrastructures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。