arXiv:2505.24615cs.CL2025-05被引 7

用大模型检测科研新意,构建了营销与NLP领域的新颖性数据集。

Harnessing Large Language Models for Scientific Novelty Detection

  • 基于论文关系提取闭包集,用大模型提炼核心思想。
  • 蒸馏大模型的思想级知识,训练轻量检索器提升准确率。
  • 在两个新数据集上表现优于现有方法,适合科研选题参考。

在科学爆炸式增长的时代,识别新颖研究想法对学术界至关重要但极具挑战。由于缺乏合适的基准数据集,相关研究受限;同时,仅依赖现有NLP技术(如检索并交叉验证)难以应对文本相似性与思想原创性之间的差距。本文提出利用大语言模型(LLMs)进行科学新颖性检测(ND),并构建了营销与自然语言处理领域的两个新数据集。为构建高质量数据集,我们基于论文间关系提取闭包集,并借助大模型总结其核心思想。为捕捉思想本质,提出通过蒸馏大模型中的思想级知识,训练轻量级检索器,实现思想层面的精准匹配。实验表明,该方法在所提基准数据集上的思想检索与新颖性检测任务中均持续优于基线方法。代码与数据已公开于https://anonymous.4open.science/r/NoveltyDetection-10FB/。

原文摘要 · Abstract (English)

In an era of exponential scientific growth, identifying novel research ideas is crucial and challenging in academia. Despite potential, the lack of an appropriate benchmark dataset hinders the research of novelty detection. More importantly, simply adopting existing NLP technologies, e.g., retrieving and then cross-checking, is not a one-size-fits-all solution due to the gap between textual similarity and idea conception. In this paper, we propose to harness large language models (LLMs) for scientific novelty detection (ND), associated with two new datasets in marketing and NLP domains. To construct the considerate datasets for ND, we propose to extract closure sets of papers based on their relationship, and then summarize their main ideas based on LLMs. To capture idea conception, we propose to train a lightweight retriever by distilling the idea-level knowledge from LLMs to align ideas with similar conception, enabling efficient and accurate idea retrieval for LLM novelty detection. Experiments show our method consistently outperforms others on the proposed benchmark datasets for idea retrieval and ND tasks. Codes and data are available at https://anonymous.4open.science/r/NoveltyDetection-10FB/.

大模型新颖性检测科研辅助知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。