用Transformer模型自动识别圣经希伯来文中的文本呼应,提升研究效率。
Intertextual Parallel Detection in Biblical Hebrew: A Transformer-Based Benchmark
- 采用E5、AlephBERT等预训练模型生成词向量,通过相似度匹配检测文本呼应。
- E5在识别呼应段落上表现最佳,AlephBERT则更擅长区分非呼应段落。
- 为古代文本研究提供自动化工具,适合语言学与神学研究者使用。
识别圣经希伯来文(BH)中的文本呼应关系是圣经学术研究的核心,传统方法依赖人工比对,耗时且易出错。本研究评估了E5、AlephBERT、MPNet和LaBSE等预训练的Transformer模型在《希伯来圣经》中检测文本呼应的潜力。聚焦《撒母耳记/列王纪》与《历代志》之间的已知呼应段落,评估各模型生成词嵌入以区分呼应与非呼应段落的能力。基于余弦相似度和沃尔德斯泰因距离(Wasserstein Distance)度量,发现E5在呼应检测上表现优异,而AlephBERT在区分非呼应段落方面更具优势。结果表明,预训练模型可显著提升古代文本中互文关系检测的效率与准确性,为古代语言研究提供新范式。
原文摘要 · Abstract (English)
Identifying parallel passages in biblical Hebrew (BH) is central to biblical scholarship for understanding intertextual relationships. Traditional methods rely on manual comparison, a labor-intensive process prone to human error. This study evaluates the potential of pre-trained transformer-based language models, including E5, AlephBERT, MPNet, and LaBSE, for detecting textual parallels in the Hebrew Bible. Focusing on known parallels between Samuel/Kings and Chronicles, I assessed each model's capability to generate word embeddings distinguishing parallel from non-parallel passages. Using cosine similarity and Wasserstein Distance measures, I found that E5 and AlephBERT show promise; E5 excels in parallel detection, while AlephBERT demonstrates stronger non-parallel differentiation. These findings indicate that pre-trained models can enhance the efficiency and accuracy of detecting intertextual parallels in ancient texts, suggesting broader applications for ancient language studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。