arXiv:2601.18085cs.HCcs.AI2026-01

用虚拟学员测试AI临床评估系统,确保评分可靠再用于真人教学。

"Crash Test Dummies" for AI-Enabled Clinical Assessment: Validating Virtual Patient Scenarios with Virtual Learners

  • 用可调能力的虚拟学员+多AI评分器模拟真实评估流程。
  • 模型准确还原虚拟学员能力,且不同评分器表现稳定一致。
  • 提出分级安全部署方案,适合医学院和AI评估系统开发者。

在医学教育中,AI用于评估临床能力,如虚拟标准化患者。但现有评估多依赖AI与人类评分者间的一致性,缺乏对案例、学习者与评分者共同影响分数的测量框架,导致结果可靠性存疑,可能误导学习者。本文通过构建带有可调ACGME标准能力维度的虚拟学员,结合多个独立的AI评分器,使用结构化关键特征项评分,并采用贝叶斯HRM-SDT模型分析转录文本。该模型将评分视为不确定性下的决策,分离学习者能力、案例表现与评分者行为,参数通过马尔可夫链蒙特卡洛(MCMC)估计。结果显示,模型能有效恢复虚拟学员的真实能力,各ACGME领域相关性显著;同时可量化各案例难度,并稳定检测出不同种子生成的AI评分器的敏感度与评分标准(严苛/宽松阈值)。此外,提出一个基于授权验证里程碑的分阶段“安全蓝图”,用于指导AI工具在真人学习者中的部署。结论表明,结合专用虚拟患者平台与严谨心理测量模型,可实现可解释、可泛化的能力评估,并支持在真实使用前完成系统验证。

原文摘要 · Abstract (English)

Background: In medical and health professions education (HPE), AI is increasingly used to assess clinical competencies, including via virtual standardized patients. However, most evaluations rely on AI-human interrater reliability and lack a measurement framework for how cases, learners, and raters jointly shape scores. This leaves robustness uncertain and can expose learners to misguidance from unvalidated systems. We address this by using AI "simulated learners" to stress-test and psychometrically characterize assessment pipelines before human use. Objective: Develop an open-source AI virtual patient platform and measurement model for robust competency evaluation across cases and rating conditions. Methods: We built a platform with virtual patients, virtual learners with tunable ACGME-aligned competency profiles, and multiple independent AI raters scoring encounters with structured Key-Features items. Transcripts were analyzed with a Bayesian HRM-SDT model that treats ratings as decisions under uncertainty and separates learner ability, case performance, and rater behavior; parameters were estimated with MCMC. Results: The model recovered simulated learners' competencies, with significant correlations to the generating competencies across all ACGME domains despite a non-deterministic pipeline. It estimated case difficulty by competency and showed stable rater detection (sensitivity) and criteria (severity/leniency thresholds) across AI raters using identical models/prompts but different seeds. We also propose a staged "safety blueprint" for deploying AI tools with learners, tied to entrustment-based validation milestones. Conclusions: Combining a purpose-built virtual patient platform with a principled psychometric model enables robust, interpretable, generalizable competency estimates and supports validation of AI-assisted assessment prior to use with human learners.

临床评估虚拟患者心理测量AI教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。