用知识图谱模拟真实医患对话,动态评估临床大模型表现
MedKGEval: A Knowledge Graph-Based Multi-Turn Evaluation Framework for Open-Ended Patient Interactions with Clinical LLMs
- 基于医学知识图谱构建患者仿真机制,实现拟人化对话行为
- 逐轮实时评估模型回复的合理性、事实正确性和安全性
- 可识别传统方法忽略的细微缺陷,适合临床AI安全测试
临床大模型的可靠评估仍面临挑战,尤其在捕捉真实医疗环境中多轮医患互动的复杂性方面。现有方法多依赖事后审查完整对话记录,忽视了对话的动态性与情境敏感性。本文提出MedKGEval,一种基于知识图谱的多轮评估框架:首先,通过整合开源资源与专家标注数据构建知识图谱,并设计控制模块从图中检索医学事实,使患者代理具备类人对话行为;其次,引入实时逐轮评估机制,由裁判智能体使用细粒度任务专用指标,在对话进程中评估模型响应的临床适宜性、事实准确性和安全性;最后,构建涵盖八种先进LLM的多轮基准测试,验证其能发现传统评估流程遗漏的细微行为缺陷与安全风险。该框架最初面向中英文医疗应用,但可通过切换知识图谱轻松扩展至其他语言,支持跨语言与领域适配。
原文摘要 · Abstract (English)
The reliable evaluation of large language models (LLMs) in medical applications remains an open challenge, particularly in capturing the complexity of multi-turn doctor-patient interactions that unfold in real clinical environments. Existing evaluation methods typically rely on post hoc review of full conversation transcripts, thereby neglecting the dynamic, context-sensitive nature of medical dialogues and the evolving informational needs of patients. In this work, we present MedKGEval, a novel multi-turn evaluation framework for clinical LLMs grounded in structured medical knowledge. Our approach introduces three key contributions: (1) a knowledge graph-driven patient simulation mechanism, where a dedicated control module retrieves relevant medical facts from a curated knowledge graph, thereby endowing the patient agent with human-like and realistic conversational behavior. This knowledge graph is constructed by integrating open-source resources with additional triples extracted from expert-annotated datasets; (2) an in-situ, turn-level evaluation framework, where each model response is assessed by a Judge Agent for clinical appropriateness, factual correctness, and safety as the dialogue progresses using a suite of fine-grained, task-specific metrics; (3) a comprehensive multi-turn benchmark of eight state-of-the-art LLMs, demonstrating MedKGEval's ability to identify subtle behavioral flaws and safety risks that are often overlooked by conventional evaluation pipelines. Although initially designed for Chinese and English medical applications, our framework can be readily extended to additional languages by switching the input knowledge graphs, ensuring seamless bilingual support and domain-specific applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。