arXiv:2604.10874cs.CLcs.AI2026-04

用检索增强提升大模型在毒理学路径分析中的准确性

AOP-Smart: A RAG-Enhanced Large Language Model Framework for Adverse Outcome Pathway Analysis

  • 基于AOP-Wiki数据构建检索增强框架,精准匹配关键事件与关系
  • 三款大模型准确率从不足35%提升至95%以上,消除幻觉问题
  • 适合毒理研究、风险评估领域,助力科学决策

不良结局通路(AOP)是毒理学研究与风险评估的重要知识框架。近年来,大语言模型(LLMs)被逐步应用于AOP相关问答与机制推理任务,但因存在幻觉问题——即生成与事实不符或缺乏证据的内容,其可靠性仍受限。为此,本研究提出面向AOP的检索增强生成(RAG)框架AOP-Smart。该方法基于AOP-Wiki官方XML数据,利用关键事件(KEs)、关键事件关系(KERs)及特定AOP信息,为用户问题检索相关知识,从而提升大模型生成结果的可靠性。为验证方法有效性,本研究构建了包含20个AOP相关问答任务的测试集,涵盖关键事件识别、上下游事件检索及复杂AOP检索任务。在Gemini、DeepSeek和ChatGPT三款主流大模型上进行对比实验,分别在无RAG与有RAG两种设置下测试。结果显示:未使用RAG时,GPT、DeepSeek与Gemini的准确率分别为15.0%、35.0%和20.0%;使用RAG后,准确率分别提升至95.0%、100.0%和95.0%。结果表明,AOP-Smart能显著缓解大模型在AOP知识任务中的幻觉问题,大幅提升答案的准确性和一致性。

原文摘要 · Abstract (English)

Adverse Outcome Pathways (AOPs) are an important knowledge framework in toxicological research and risk assessment. In recent years, large language models (LLMs) have gradually been applied to AOP-related question answering and mechanistic reasoning tasks. However, due to the existence of the hallucination problem, that is, the model may generate content that is inconsistent with facts or lacks evidence, their reliability is still limited. To address this issue, this study proposes an AOP-oriented Retrieval-Augmented Generation (RAG) framework, AOP-Smart. Based on the official XML data from AOP-Wiki, this method uses Key Events (KEs), Key Event Relationships (KERs), and specific AOP information to retrieve relevant knowledge for user questions, thereby improving the reliability of the generated results of large language models. To evaluate the effectiveness of the proposed method, this study constructed a test set containing 20 AOP-related question answering tasks, covering KE identification, upstream and downstream KE retrieval, and complex AOP retrieval tasks. Experiments were conducted on three mainstream large language models, Gemini, DeepSeek, and ChatGPT, and comparative tests were performed under two settings: without RAG and with RAG. The experimental results show that, without using RAG, the accuracies of GPT, DeepSeek, and Gemini were 15.0\%, 35.0\%, and 20.0\%, respectively; after using RAG, their accuracies increased to 95.0\%, 100.0\%, and 95.0\%, respectively. The results indicate that AOP-Smart can significantly alleviate the hallucination problem of large language models in AOP knowledge tasks, and greatly improve the accuracy and consistency of their answers.

毒理学RAG大模型知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。