arXiv:2603.14674cs.CL2026-03

用语义相似度分析梅尔维尔阅读与写作的关系

Computational Analysis of Semantic Connections Between Herman Melville Reading and Writing

  • 用BERTScore比对梅尔维尔作品与藏书文本的语义相似性
  • 准确识别出专家确认的相似段落,发现新潜在影响痕迹
  • 为文学影响研究提供可量化的计算支持方法

本研究通过计算语义相似性分析,探讨赫尔曼·梅尔维尔阅读对其创作的影响。基于其实际拥有或阅读过的书籍记录,将梅尔维尔作品中的选段与图书馆文本进行对比。方法包括句子级和非重叠5-词组级别的文本分割,使用BERTScore计算相似度。不采用固定阈值判断引用,而是将精确率、召回率和F1分数视为可能语义对齐的指示,暗示文学影响的存在。实验结果表明,该方法成功捕捉到专家确认的相似实例,并揭示了需进一步质性考察的其他段落。研究结果表明,语义相似性方法为文学源流与影响研究提供了有效的计算框架。

原文摘要 · Abstract (English)

This study investigates the potential influence of Herman Melville reading on his own writings through computational semantic similarity analysis. Using documented records of books known to have been owned or read by Melville, we compare selected passages from his works with texts from his library. The methodology involves segmenting texts at both sentence level and non-overlapping 5-gram level, followed by similarity computation using BERTScore. Rather than applying fixed thresholds to determine reuse, we interpret precision, recall, and F1 scores as indicators of possible semantic alignment that may suggest literary influence. Experimental results demonstrate that the approach successfully captures expert-identified instances of similarity and highlights additional passages warranting further qualitative examination. The findings suggest that semantic similarity methods provide a useful computational framework for supporting source and influence studies in literary scholarship.

文学分析语义相似计算人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。