arXiv:2502.20541cs.CLcs.IR2025-02被引 8

用智能检索增强生成系统,帮纳米科技研究者高效找文献。

NANOGPT: A Query-Driven Large Language Model Retrieval-Augmented Generation System for Nanotechnology Research

  • 基于多源检索的查询驱动架构,整合谷歌学术与多家出版商数据。
  • 测试显示文献检索效率大幅提升,准确率和相关性优于通用大模型。
  • 适合纳米科技领域研究人员快速完成综述与文献调研。

本文提出一种专为纳米科技研究设计的大型语言模型检索增强生成(LLM-RAG)系统。该系统利用先进语言模型作为智能研究助手,提升纳米科技领域文献综述的效率与全面性。核心是先进的查询后端检索机制,整合多个权威数据源:通过谷歌学术高级搜索获取文献,并从爱思唯尔、施普林格自然及美国化学会出版社的开放获取论文中抓取数据。这种多源策略确保了最新、多样且广泛的学术文章覆盖。系统经过严格测试,验证了其在显著缩短文献调研时间与精力消耗的同时,保持高准确性与查询相关性,性能优于标准公开大模型,展现出加速纳米科技研究进展的巨大潜力。

原文摘要 · Abstract (English)

This paper presents the development and application of a Large Language Model Retrieval-Augmented Generation (LLM-RAG) system tailored for nanotechnology research. The system leverages the capabilities of a sophisticated language model to serve as an intelligent research assistant, enhancing the efficiency and comprehensiveness of literature reviews in the nanotechnology domain. Central to this LLM-RAG system is its advanced query backend retrieval mechanism, which integrates data from multiple reputable sources. The system retrieves relevant literature by utilizing Google Scholar's advanced search, and scraping open-access papers from Elsevier, Springer Nature, and ACS Publications. This multifaceted approach ensures a broad and diverse collection of up-to-date scholarly articles and papers. The proposed system demonstrates significant potential in aiding researchers by providing a streamlined, accurate, and exhaustive literature retrieval process, thereby accelerating research advancements in nanotechnology. The effectiveness of the LLM-RAG system is validated through rigorous testing, illustrating its capability to significantly reduce the time and effort required for comprehensive literature reviews, while maintaining high accuracy, query relevance and outperforming standard, publicly available LLMS.

大模型文献检索纳米科技RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。