arXiv:2509.15292cs.AI2025-09被引 2

用语义相似度自动找相关论文,省时高效。

An Artificial Intelligence Driven Semantic Similarity-Based Pipeline for Rapid Literature

  • 基于Transformer嵌入和余弦相似度匹配论文
  • 三类模型评估,统计阈值过滤有效文献
  • 无需人工标注,适合快速探索新领域

我们提出一种基于语义相似度的自动化文献综述流程。与传统系统或优化方法不同,该方法强调低开销与高相关性,利用基于Transformer的嵌入和余弦相似度实现。输入论文标题和摘要后,系统生成关键词,从开放获取库中检索相关论文,并根据语义接近程度排序。评估了三种嵌入模型,采用统计阈值法过滤相关论文,构建高效文献综述流程。尽管未使用启发式反馈或真实相关性标签,该系统在初步研究与探索性分析中展现出可扩展性和实用性。

原文摘要 · Abstract (English)

We propose an automated pipeline for performing literature reviews using semantic similarity. Unlike traditional systematic review systems or optimization based methods, this work emphasizes minimal overhead and high relevance by using transformer based embeddings and cosine similarity. By providing a paper title and abstract, it generates relevant keywords, fetches relevant papers from open access repository, and ranks them based on their semantic closeness to the input. Three embedding models were evaluated. A statistical thresholding approach is then applied to filter relevant papers, enabling an effective literature review pipeline. Despite the absence of heuristic feedback or ground truth relevance labels, the proposed system shows promise as a scalable and practical tool for preliminary research and exploratory analysis.

文献综述语义相似度自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。