构建患者仿真框架,评估对话式医疗AI在不同患者特征下的风险表现
A Patient Simulation Framework for Risk Assessment of Conversational Healthcare AI: Evaluation of an Antidepressant Decision Aid
- 基于医疗、语言和行为三维度构建患者仿真模型
- 健康素养越低,AI推荐准确率越差,概念召回率从47.6%降至81.9%
- 符合NIST AI风险框架,适合医疗AI安全评估与验证
本研究开发并验证了一种符合美国国家标准与技术研究院(NIST)人工智能风险管理框架(AI RMF)MAP和MEASURE功能的患者仿真框架,为对话式临床AI在医学、语言和行为层面的患者差异中识别与刻画性能风险提供实证基础。我们将其应用于重度抑郁症抗抑郁药选择的对话式决策辅助系统(AI Decision Aid)。模拟器整合三个维度:(1)基于All of Us电子健康记录的风险比筛选构建的医疗特征;(2)反映健康素养梯度和疾病特异性沟通的语言特征;(3)合作、分心与对抗性交互的行为特征。生成500次模拟对话,通过人工标注和大模型判断评估特征保真度,并检验对AI Decision Aid概念检索与药物推荐的下游影响。结果表明,医疗概念表达保真度高(8,210个概念中准确率达96.6%),人工标注者间一致性κ=0.73,大模型判断与人工一致κ=0.78;行为特征区分可靠(κ=0.93),语言特征中等一致(κ=0.61)。框架揭示了随着健康素养下降,AI决策辅助性能单调退化:概念检索的Rank-1准确率从有限健康素养者的47.6%升至熟练者81.9%,而抗抑郁药推荐准确率相应下降。
原文摘要 · Abstract (English)
Objective: This study develops and validates a patient simulation framework that aligns with the National Institute of Standards and Technology (NIST) AI Risk Management Framework (AI RMF) MAP and MEASURE functions, providing an empirical basis for identifying and characterizing performance risks in conversational clinical AI across medical, linguistic, and behavioral patient variation. We applied the framework to a conversational decision aid for antidepressant selection in major depressive disorder (the AI Decision Aid). Methods: The simulator integrates three profile dimensions: (1) medical profiles constructed from All of Us electronic health records using risk-ratio gating; (2) linguistic profiles modeling a health literacy gradient and condition-specific communication; and (3) behavioral profiles representing cooperative, distracted, and adversarial engagement. We generated 500 simulated conversations and evaluated profile fidelity through human annotation and an LLM judge, then assessed downstream effects on the AI Decision Aid's concept retrieval and antidepressant recommendations. Results: The patient simulator expressed medical concepts with high fidelity (96.6% accurate across 8,210 concepts), with human inter-annotator agreement of 0.73 $κ$ and LLM-judge agreement against human annotators of 0.78 $κ$. Behavioral profiles were reliably distinguished (0.93 $κ$), and linguistic profiles showed moderate agreement (0.61 $κ$). The framework revealed monotonic degradation in AI Decision Aid performance across the health literacy gradient. Rank-1 concept retrieval increased from 47.6% for limited health literacy to 81.9% for proficient health literacy, with corresponding declines in antidepressant recommendation accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。