用大模型提升科学文献主题发现,更准更快。
SciTopic: Enhancing Topic Discovery in Scientific Literature through Advanced LLM
- 基于大模型构建文本编码器,融合标题摘要等信息。
- 通过熵采样与三元组优化,增强主题相关性识别。
- 在三个数据集上超越现有方法,适合科研趋势分析。
科学文献的主题发现为研究人员识别新兴趋势、探索新研究方向提供了重要支持,有助于科学信息检索。尽管已有大量机器学习方法,尤其是深度嵌入技术被用于主题发现,但多数方法依赖词向量捕捉语义,缺乏对科学出版物的全面理解,难以处理复杂高维文本关系。受大语言模型(LLMs)出色文本理解能力启发,我们提出一种由大模型增强的先进主题发现方法——SciTopic。首先,构建文本编码器以捕获科学文献的内容,包括元数据、标题和摘要;其次,设计空间优化模块,结合基于熵的采样与由大模型指导的三元组任务,强化对主题相关性和模糊实例间上下文细节的关注;然后,基于大模型引导,通过优化三元组对比损失对文本编码器进行微调,促使编码器更好区分不同主题的实例;最后,在三个真实世界科学文献数据集上的广泛实验表明,SciTopic显著优于当前最优(SOTA)的科学主题发现方法,使研究人员能够获得更深入、更快速的洞察。
原文摘要 · Abstract (English)
Topic discovery in scientific literature provides valuable insights for researchers to identify emerging trends and explore new avenues for investigation, facilitating easier scientific information retrieval. Many machine learning methods, particularly deep embedding techniques, have been applied to discover research topics. However, most existing topic discovery methods rely on word embedding to capture the semantics and lack a comprehensive understanding of scientific publications, struggling with complex, high-dimensional text relationships. Inspired by the exceptional comprehension of textual information by large language models (LLMs), we propose an advanced topic discovery method enhanced by LLMs to improve scientific topic identification, namely SciTopic. Specifically, we first build a textual encoder to capture the content from scientific publications, including metadata, title, and abstract. Next, we construct a space optimization module that integrates entropy-based sampling and triplet tasks guided by LLMs, enhancing the focus on thematic relevance and contextual intricacies between ambiguous instances. Then, we propose to fine-tune the textual encoder based on the guidance from the LLMs by optimizing the contrastive loss of the triplets, forcing the text encoder to better discriminate instances of different topics. Finally, extensive experiments conducted on three real-world datasets of scientific publications demonstrate that SciTopic outperforms the state-of-the-art (SOTA) scientific topic discovery methods, enabling researchers to gain deeper and faster insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。