测试9种RAG配置,发现用OpenAI模型答对率高达91.4%。
Evaluating Retrieval-Augmented Generation Agents for Autonomous Scientific Discovery in Astrophysics
- 用人类专家评估945条生成答案,筛选最优RAG配置
- 最佳配置在宇宙学问答上达到91.4%准确率
- 构建可扩展的LLM评标准则,助力自动化科研
我们针对宇宙学领域专门构建了105个问答对,评估了9种检索增强生成(RAG)代理配置。每种配置由人类专家进行人工评估,共评估945条生成回答。结果显示,使用OpenAI嵌入和生成模型的RAG配置表现最佳,准确率达91.4%。基于人工评估结果,我们校准了大语言模型作为裁判(LLMaaJ)系统,可作为人类评估的可靠替代方案。该成果使我们能系统性选择适用于天体物理自主科学发现多代理系统的最优RAG配置,并提供一个可扩展至数千个宇宙学问答对的评估工具。我们公开了问答数据集、人工评估结果、RAG流水线及LLMaaJ系统,供天体物理学界进一步使用。
原文摘要 · Abstract (English)
We evaluate 9 Retrieval Augmented Generation (RAG) agent configurations on 105 Cosmology Question-Answer (QA) pairs that we built specifically for this purpose.The RAG configurations are manually evaluated by a human expert, that is, a total of 945 generated answers were assessed. We find that currently the best RAG agent configuration is with OpenAI embedding and generative model, yielding 91.4\% accuracy. Using our human evaluation results we calibrate LLM-as-a-Judge (LLMaaJ) system which can be used as a robust proxy for human evaluation. These results allow us to systematically select the best RAG agent configuration for multi-agent system for autonomous scientific discovery in astrophysics (e.g., cmbagent presented in a companion paper) and provide us with an LLMaaJ system that can be scaled to thousands of cosmology QA pairs. We make our QA dataset, human evaluation results, RAG pipelines, and LLMaaJ system publicly available for further use by the astrophysics community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。