用n-gram嵌入量化文本间引用关系,实现可扩展的文学关联分析。
Modelling Intertextuality with N-gram Embeddings
- 通过对比两文本n-gram嵌入的相似度,平均得整体互文性得分。
- 在4个已知互文强度的文本上验证有效,267个文本测试展现高效性。
- 可揭示文本网络中的核心节点与社区结构,适合文学分析与数字人文研究。
互文性是文学研究的核心概念,指文本间通过各类引用建立的复杂联系。本文提出一种新的定量互文性模型,通过计算两文本n-gram嵌入的成对相似性并取均值,实现可扩展的分析与网络洞察。在4个已知互文程度的文本上进行验证,并对267个多样化文本进行可扩展性测试,结果表明该方法高效且有效。网络分析进一步揭示了中心性与社群结构,证实其在捕捉和量化互文关系方面的成功。
原文摘要 · Abstract (English)
Intertextuality is a central tenet in literary studies. It refers to the intricate links between literary texts that are created by various types of references. This paper proposes a new quantitative model of intertextuality to enable scalable analysis and network-based insights: perform pairwise comparisons of the embeddings of n-grams from two texts and average their results as the overall intertextuality. Validation on four texts with known degrees of intertextuality, alongside a scalability test on 267 diverse texts, demonstrates the method's effectiveness and efficiency. Network analysis further reveals centrality and community structures, affirming the approach's success in capturing and quantifying intertextual relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。