arXiv:2601.07533cs.IRcs.CL2026-01被引 3

构建首个拉丁文文本互文性检测基准,助力学者追溯古代文献影响脉络。

Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

  • 构建包含17.6万段落和1490组专家验证互文关系的标注数据集
  • 在945个已知引用上建立检索与分类基线,模型表现优于传统词汇方法
  • 适合古典学、数字人文及跨语言语义匹配研究者使用

追踪历史文本间的关联是互文研究的重要环节,有助于重构作者的虚拟书库并识别其创作受哪些文本影响。这些互文关系形式多样,从直接引文到经过形态变化伪装的隐喻和转述皆有。语言模型因其超越词汇重叠捕捉语义相似性的能力,为该任务提供了新路径。然而,新方法的发展受限于缺乏标准化基准与易用数据集。为此,本文提出Loci Similes,一个面向拉丁文互文性检测的基准,包含约17.6万段文本和1,490组经专家验证的互文关系,其中包含945个来自现有数据集的标注引用。基于此数据,我们为互文性检索与分类建立了词法方法和预训练编码器语言模型的基线性能。

原文摘要 · Abstract (English)

Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual links manifest in diverse forms, ranging from direct verbatim quotations to subtle allusions and paraphrases disguised by morphological variation. Language models offer a promising path forward due to their capability of capturing semantic similarity beyond lexical overlap. However, the development of new methods for this task is held back by the scarcity of standardized benchmarks and easy-to-use datasets. We address this gap by introducing Loci Similes, a benchmark for Latin intertextuality detection comprising a curated dataset of ~176k text segments and 1,490 expert-verified parallels, including 945 labeled references from an existing dataset. Using this data, we establish baselines for retrieval and classification of intertextualities with both lexical methods and pretrained encoder language models.

互文性拉丁文语言模型数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。