用开源模型搭建科学文献问答系统,保护科研数据隐私
Retrieval-Augmented Question Answering over Scientific Literature for the Electron-Ion Collider
- 基于arXiv的EIC文献构建本地知识库,结合LLaMA生成答案
- 无需上传未发表数据,实现安全高效的领域问答
- 适合实验核物理领域研究人员快速获取专业信息
为提升语言模型在回答特定领域技术问题中的能力,本工作开发了一款受检索增强生成(RAG)启发的问答应用。该系统基于自建数据库,索引了与电子-离子对撞机(EIC)实验相关的arXiv论文,并采用开源LLaMA模型生成答案。这是此前基于专有模型和云端外部知识库系统的延伸,现改为本地部署的RAG系统,提供成本低、资源受限环境下的可行替代方案。该架构保障数据隐私,避免将预发布科学数据发送至公共领域。未来将扩展知识库至异构的EIC相关出版物与报告,并将应用流程升级至LangGraph框架。
原文摘要 · Abstract (English)
To harness the power of Language Models in answering domain specific specialized technical questions, Retrieval Augmented Generation (RAG) is been used widely. In this work, we have developed a Q\&A application inspired by the Retrieval Augmented Generation (RAG), which is comprised of an in-house database indexed on the arXiv articles related to the Electron-Ion Collider (EIC) experiment - one of the largest international scientific collaboration and incorporated an open-source LLaMA model for answer generation. This is an extension to it's proceeding application built on proprietary model and Cloud-hosted external knowledge-base for the EIC experiment. This locally-deployed RAG-system offers a cost-effective, resource-constraint alternative solution to build a RAG-assisted Q\&A application on answering domain-specific queries in the field of experimental nuclear physics. This set-up facilitates data-privacy, avoids sending any pre-publication scientific data and information to public domain. Future improvement will expand the knowledge base to encompass heterogeneous EIC-related publications and reports and upgrade the application pipeline orchestration to the LangGraph framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。