构建临床对话安全评估框架,用多智能体模拟真实医疗场景
MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
- 设计多智能体仿真框架,整合安全分类与自动化评估工具
- 行为判别器在240轮对话中达0.96的F1值,超人工专家
- 支持监管级安全审计,适合医疗AI研发与合规验证
尽管大语言模型在临床对话系统中应用日益广泛,现有评估仍局限于任务完成度或流畅性,难以反映安全关键系统的行为规范与风险管控需求。本文提出MATRIX(多智能体仿真框架,用于安全交互与情境化临床对话评估),一个结构化、可扩展的安全导向评估框架。该框架包含三部分:(1) 基于结构化安全工程方法构建的临床场景、预期行为与失效模式分类体系;(2) BehvJudge,一个基于LLM的评估器,用于检测对话中的安全相关失败,经临床专家标注验证;(3) PatBot,一个模拟患者代理,能生成多样化、场景适配的回应,通过人因学专家评估其真实感与行为保真度,并开展患者偏好研究。在三个实验中,我们证明MATRIX可实现系统化、可扩展的安全评估。BehvJudge使用Gemini 2.5-Pro,在240轮盲评对话中达到0.96的F1值与0.999的灵敏度,表现优于临床医生。我们还开展了首个针对基于LLM患者模拟的真实感分析,证实PatBot在定量与定性评估中均能可靠模拟真实患者行为。利用MATRIX,我们在2100轮模拟对话中对五个LLM代理在14种风险场景和10个临床领域进行了基准测试。MATRIX是首个将结构化安全工程与可扩展、可验证的对话式AI评估统一的框架,支持监管对齐的安全审计。所有评估工具、提示、结构化场景及数据集均已公开。
原文摘要 · Abstract (English)
Despite the growing use of large language models (LLMs) in clinical dialogue systems, existing evaluations focus on task completion or fluency, offering little insight into the behavioral and risk management requirements essential for safety-critical systems. This paper presents MATRIX (Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation), a structured, extensible framework for safety-oriented evaluation of clinical dialogue agents. MATRIX integrates three components: (1) a safety-aligned taxonomy of clinical scenarios, expected system behaviors and failure modes derived through structured safety engineering methods; (2) BehvJudge, an LLM-based evaluator for detecting safety-relevant dialogue failures, validated against expert clinician annotations; and (3) PatBot, a simulated patient agent capable of producing diverse, scenario-conditioned responses, evaluated for realism and behavioral fidelity with human factors expertise, and a patient-preference study. Across three experiments, we show that MATRIX enables systematic, scalable safety evaluation. BehvJudge with Gemini 2.5-Pro achieves expert-level hazard detection (F1 0.96, sensitivity 0.999), outperforming clinicians in a blinded assessment of 240 dialogues. We also conducted one of the first realism analyses of LLM-based patient simulation, showing that PatBot reliably simulates realistic patient behavior in quantitative and qualitative evaluations. Using MATRIX, we demonstrate its effectiveness in benchmarking five LLM agents across 2,100 simulated dialogues spanning 14 hazard scenarios and 10 clinical domains. MATRIX is the first framework to unify structured safety engineering with scalable, validated conversational AI evaluation, enabling regulator-aligned safety auditing. We release all evaluation tools, prompts, structured scenarios, and datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。