用AI识别麦卡锡小说中的圣经典故,发现349处隐晦引用。
'The Order in the Horse's Heart': A Case Study in LLM-Assisted Stylometry for the Discovery of Biblical Allusion in Modern Literary Fiction
- 分上下两路检测:词汇罕见性+上下文嵌入,结合大模型判断
- 识别出349处文学典故,对已知典故的召回率达54%
- 适合研究文学互文性与大模型辅助文本分析的学者
本文提出一种双轨式管道,用于检测现代文学小说中的圣经典故,并应用于科马克·麦卡锡的作品。底层嵌入路径利用逆文档频率识别与《钦定版圣经》共享的稀有词汇,将出现位置嵌入局部语境以消歧义,并通过级联大语言模型审查候选段落对;顶层注册路径让大模型无导向地阅读麦卡锡文风,捕捉非词汇罕见性特征的典故。两条路径由长上下文模型交叉验证,将整部小说与KJV置于同一处理流中,并逐项对照学术文献。限定于具有文字回响(共用措辞、重构词汇或移植节奏)的典故,区分文学典故与显性引用(如比喻提及圣经人物),共发现349处典故。在115个已有文献记载的典故中,该管道独立恢复62处(54%召回率),按类型从30%(意象转化)到80%(风格重叠)不等。结果说明大模型在增强机械风格分析上的价值,及其在大规模文学语料中开展互文性统计研究的潜力。
原文摘要 · Abstract (English)
We present a dual-track pipeline for detecting biblical allusions in literary fiction and apply it to the novels of Cormac McCarthy. A bottom-up embedding track uses inverse document frequency to identify rare vocabulary shared with the King James Bible, embeds occurrences in their local context for sense disambiguation, and passes candidate passage pairs through cascaded LLM review. A top-down register track asks an LLM to read McCarthy's prose undirected to any specific biblical passage for comparison, catching allusions not distinguished by word or phrase rarity. Both tracks are cross-validated by a long-context model that holds entire novels alongside the KJV in a single pass, and every finding is checked against published scholarship. Restricting attention to allusions that carry a textual echo--shared phrasing, reworked vocabulary, or transplanted cadence--and distinguishing literary allusions proper from signposted biblical references (similes naming biblical figures, characters overtly citing scripture), the pipeline surfaces 349 allusions across the corpus. Among a target set of 115 previously documented allusions retrieved through human review of the academic literature, the pipeline independently recovers 62 (54% recall), with recall varying by connection type from 30% (transformed imagery) to 80% (register collisions). We contextualise these results with respect to the value-add from LLMs as assistants to mechanical stylometric analyses, and their potential to facilitate the statistical study of intertextuality in massive literary corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。