用多模态数据验证虚拟医考中考官判断,提升评分公正性
Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs

- 通过视频、日志和标注比对考官陈述与实际操作
- 错误检测达76.7%召回率,纠正后正确率从39.2%升至79.2%
- 适合医疗评估系统优化与客观化考试研究者
客观结构化临床检查(OSCE)是评估临床能力的金标准,但评分易受考官主观性、疲劳和认知偏见影响。传统考官间一致性统计缺乏对错误根源的解释力,既不分析考官推理过程,也无法验证其陈述是否符合实际事件。为此,我们提出质量行为保障(QAA)框架,通过对比虚拟现实(VR)儿科OSCE中考官声称的行为与由视频、VR日志及演员标注构建的参考事件记录,实现多模态验证。QAA结合约束时间动作对齐模型(实现动作定位与行为源归属)与大语言模型(提取考官陈述并比对记录)。在五折交叉验证中,动作识别F1达到99.2%±0.7,时间对齐W@16为93.4%±1.9。整体上,QAA以69.9%精确率、76.7%召回率检测考官错误;回溯评估显示,纠正这些错误使真实正确的转录比例从39.2%提升至79.2%,支持更公平的OSCE质量评估。
原文摘要 · Abstract (English)
Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistics lacks explanatory power regarding the source of errors, as it neither analyzes examiner reasoning nor verifies examiner claims against actual events. Thus, we introduce Quality Action Assurance (QAA), a multimodal framework that verifies examiner claims in Virtual Reality (VR) pediatric OSCEs by comparing actions claimed by examiners against a reference record of events constructed from video, VR logs, and actor annotations. QAA combines a constrained temporal action alignment model, which performs action localization and actor source attribution, with a large language model that extracts examiner claims and checks them against the record. Across a 5-fold cross-validation, QAA achieves 99.2\% $\pm$ 0.7\% Actor F1 and 93.4\% $\pm$ 1.9\% W@16 for temporal alignment. Overall, QAA detects examiner errors with 69.9\% precision and 76.7\% recall; in retrospective evaluation, correcting the detected errors raises the share of factually correct transcripts from 39.2\% to 79.2\%, supporting fairer OSCE quality assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。