arXiv:2605.30400cs.CL2026-05被引 1

用RAG+多模型投票评估ChatGPT生成生物医学关联的可靠性

Protocol for evaluating ChatGPT in biomedical association generation and verification using a RAG-enabled, cross-model majority voting workflow

  • 构建RAG增强的跨模型投票流程,提升生成结果可信度
  • 通过文献和本体验证,发现高比例关联存在幻觉现象
  • 适合医学知识图谱构建与LLM可信性研究者参考

我们提出一种评估ChatGPT生成疾病导向生物医学关联能力的协议。该协议定义了关联生成、生物实体通过生物本体验证,以及基于文献的关联验证流程。引入自一致性策略,评估不同ChatGPT模型间的生成可靠性。为克服本体精确匹配的局限,提供一个使用开源大语言模型驱动的检索增强生成(RAG)工作流,实现对其他大语言模型生成内容的语义验证,揭示其幻觉问题。

原文摘要 · Abstract (English)

We present a protocol to evaluate ChatGPT's ability to generate disease-centric biomedical associations. It outlines how we generate the associations, validate the biological entities using biomedical ontologies, and verify associations using literature. The protocol includes a self-consistency strategy to assess generative reliability across ChatGPT models. To address ontology exact-match limitations, we provide a use case performing semantic verification through a workflow enabled by Retrieval-Augmented Generation (RAG) powered by open-source large language models (LLMs). This enables LLMs to establish truth over content generated by other LLMs and expose hallucination.

大模型评估RAG生物医学幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。