用知识图谱生成迷惑性选项,让医学问答题更难考出真水平。
Enhancing Clinical Multiple-Choice Questions Benchmarks with Knowledge Graph Guided Distractor Generation
- 基于医学知识图谱的多步语义搜索,找出发病机制相似但错误的干扰项路径。
- 在6个主流医学问答数据集上,显著降低顶尖大模型准确率。
- 适合评估医疗大模型真实推理能力,尤其用于临床决策场景测试。
临床诊断与治疗依赖强决策能力,亟需严格评估基准以检验大语言模型(LLMs)的可靠性。本文提出一种基于知识图谱的数据增强框架,通过生成与正确答案语义接近但错误的干扰项(distractors),提升医学多选题(MCQ)数据集难度。利用医学知识图谱进行多步、语义导向的路径搜索,识别出医学相关但事实错误的关联路径,并引导大模型生成更具迷惑性的干扰项。该知识图谱引导的干扰项生成(KGGDG)流程应用于六个广泛使用的医学问答基准,均显著降低当前先进大模型的准确率。结果表明,KGGDG是实现更稳健、更具诊断意义的医学大模型评估的强大工具。
原文摘要 · Abstract (English)
Clinical tasks such as diagnosis and treatment require strong decision-making abilities, highlighting the importance of rigorous evaluation benchmarks to assess the reliability of large language models (LLMs). In this work, we introduce a knowledge-guided data augmentation framework that enhances the difficulty of clinical multiple-choice question (MCQ) datasets by generating distractors (i.e., incorrect choices that are similar to the correct one and may confuse existing LLMs). Using our KG-based pipeline, the generated choices are both clinically plausible and deliberately misleading. Our approach involves multi-step, semantically informed walks on a medical knowledge graph to identify distractor paths-associations that are medically relevant but factually incorrect-which then guide the LLM in crafting more deceptive distractors. We apply the designed knowledge graph guided distractor generation (KGGDG) pipline, to six widely used medical QA benchmarks and show that it consistently reduces the accuracy of state-of-the-art LLMs. These findings establish KGGDG as a powerful tool for enabling more robust and diagnostic evaluations of medical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。