挖掘文本单向引用关系,助力文学历史研究
Mining Asymmetric Intertextuality
- 分-归一-合并三步法,用大模型提取结构化元数据
- 可识别直接引用、改写及跨文档影响等多类关系
- 适合动态增长的文献库,如数字人文研究场景
本文提出自然语言处理与数字人文领域的新任务:挖掘非对称互文性。非对称互文指一个文本引用或借鉴另一个文本但不被反向引用,常见于后世作品对经典文本的引用。我们提出一种可扩展、自适应的方法,采用分-归一-合并范式:将文档切分为小块,利用大模型辅助提取元数据进行结构化归一化,并在查询时合并以检测显性和隐性互文关系。该方法可覆盖从直接引用到改写、跨文档影响等多种层级,结合元数据过滤、向量相似度搜索与大模型验证实现。系统适用于持续增长的语料库,如不断扩充的文学档案或历史数据库,支持新文档的持续集成,对文学研究、历史分析等数字人文实践具有重要价值。
原文摘要 · Abstract (English)
This paper introduces a new task in Natural Language Processing (NLP) and Digital Humanities (DH): Mining Asymmetric Intertextuality. Asymmetric intertextuality refers to one-sided relationships between texts, where one text cites, quotes, or borrows from another without reciprocation. These relationships are common in literature and historical texts, where a later work references aclassical or older text that remain static. We propose a scalable and adaptive approach for mining asymmetric intertextuality, leveraging a split-normalize-merge paradigm. In this approach, documents are split into smaller chunks, normalized into structured data using LLM-assisted metadata extraction, and merged during querying to detect both explicit and implicit intertextual relationships. Our system handles intertextuality at various levels, from direct quotations to paraphrasing and cross-document influence, using a combination of metadata filtering, vector similarity search, and LLM-based verification. This method is particularly well-suited for dynamically growing corpora, such as expanding literary archives or historical databases. By enabling the continuous integration of new documents, the system can scale efficiently, making it highly valuable for digital humanities practitioners in literacy studies, historical research and related fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。