arXiv:2512.10195cs.CLcs.LG2025-12被引 3

用虚拟病人模拟真实问诊,自动评估医疗对话模型表现

AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding

  • 将静态问答数据转为虚拟病人,构建多轮互动问诊场景
  • 提出CARE多维度指标,涵盖准确性、效率、共情与鲁棒性
  • 专家验证有效,适合开发医疗对话AI的团队参考

大语言模型在医疗领域的安全可信应用亟需有效评估。尽管已有多种静态医疗问答基准,但对动态交互式多轮临床对话的评估仍不充分,且缺乏超越简单准确率的多维度评价方法。由于患者状态与交互路径组合空间巨大,标准化定量评估困难。为此,我们提出AutoMedic——一个基于多智能体的自动化评估框架,可将现成的静态QA数据集转化为虚拟患者档案,实现真实、临床基础的多轮对话模拟。通过CARE指标(临床对话准确性、效率/策略、共情与鲁棒性)评估不同临床对话模型表现。经专家验证,AutoMedic具备有效性,为医疗对话大模型的开发提供实用指导。

原文摘要 · Abstract (English)

Evaluating large language models (LLMs) has recently emerged as a critical issue for safe and trustworthy application of LLMs in the medical domain. Although a variety of static medical question-answering (QA) benchmarks have been proposed, many aspects remain underexplored, such as the effectiveness of LLMs in generating responses in dynamic, interactive clinical multi-turn conversation situations and the identification of multi-faceted evaluation strategies beyond simple accuracy. However, formally evaluating a dynamic, interactive clinical situation is hindered by its vast combinatorial space of possible patient states and interaction trajectories, making it difficult to standardize and quantitatively measure such scenarios. Here, we introduce AutoMedic, a multi-agent simulation framework that enables automated evaluation of LLMs as clinical conversational agents. AutoMedic transforms off-the-shelf static QA datasets into virtual patient profiles, enabling realistic and clinically grounded multi-turn clinical dialogues between LLM agents. The performance of various clinical conversational agents is then assessed based on our CARE metric, which provides a multi-faceted evaluation standard of clinical conversational accuracy, efficiency/strategy, empathy, and robustness. Our findings, validated by human experts, demonstrate the validity of AutoMedic as an automated evaluation framework for clinical conversational agents, offering practical guidelines for the effective development of LLMs in conversational medical applications.

医疗对话评估框架多轮对话大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。