arXiv:2410.01141cs.CLcs.AI2024-10被引 1

用NLP和大模型评估经济论文标题去重效果,发现重复率可能很低。

Evaluating Deduplication Techniques for Economic Research Paper Titles with a Focus on Semantic Similarity using NLP and LLMs

  • 结合多种配对方法与语义相似度模型检测标题重复
  • 不同方法下语义相似度低,表明重复标题较少
  • 结果支持大模型在文本去重中的有效性,适合数据清洗场景

本研究针对大规模经济研究论文标题的NLP数据集,探索高效的去重技术。我们比较了多种配对方法,结合传统的距离度量(如编辑距离、余弦相似度)以及sBERT模型进行语义评估。结果显示,不同方法下的语义相似度普遍较低,暗示重复标题可能存在率较低。进一步通过人工标注的基准数据集验证,得出更可靠的结论。结果与基于NLP和大语言模型的距离度量发现一致。

原文摘要 · Abstract (English)

This study investigates efficient deduplication techniques for a large NLP dataset of economic research paper titles. We explore various pairing methods alongside established distance measures (Levenshtein distance, cosine similarity) and a sBERT model for semantic evaluation. Our findings suggest a potentially low prevalence of duplicates based on the observed semantic similarity across different methods. Further exploration with a human-annotated ground truth set is completed for a more conclusive assessment. The result supports findings from the NLP, LLM based distance metrics.

去重语义相似度NLP大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。