arXiv:2412.17552cs.CLcs.IR2024-12被引 2

对比三种方法在莎士比亚诗与泰勒歌词中的相似度评分表现

Comparative Analysis of Document-Level Embedding Methods for Similarity Scoring on Shakespeare Sonnets and Taylor Swift Lyrics

  • 用余弦相似度比较TF-IDF、Word2Vec均值和BERT的文本相似度计算效果
  • Word2Vec在跨领域比较中语义泛化能力更强,TF-IDF依赖词汇重叠
  • BERT表现较差,可能因缺乏特定领域微调,适合关注领域适配的研究者

本研究评估了TF-IDF加权、平均Word2Vec嵌入和BERT嵌入在两个截然不同文本领域(莎士比亚十四行诗与泰勒·斯威夫特歌词)中进行文档相似度评分的表现。通过分析余弦相似度得分,凸显了各方法的优势与局限。结果表明,TF-IDF高度依赖词汇重叠,而Word2Vec在跨领域比较中展现出更强的语义泛化能力;BERT在挑战性领域表现较低,可能源于缺乏领域特定微调。

原文摘要 · Abstract (English)

This study evaluates the performance of TF-IDF weighting, averaged Word2Vec embeddings, and BERT embeddings for document similarity scoring across two contrasting textual domains. By analysing cosine similarity scores, the methods' strengths and limitations are highlighted. The findings underscore TF-IDF's reliance on lexical overlap and Word2Vec's superior semantic generalisation, particularly in cross-domain comparisons. BERT demonstrates lower performance in challenging domains, likely due to insufficient domainspecific fine-tuning.

文本相似度嵌入方法跨域分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。