arXiv:2605.20233cs.CVcs.AI2026-05中稿 · CVPR

用视觉模型分析护士模拟训练视频,发现越熟练的学生动作越难识别,但更规范。

AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education

论文配图:AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education
图 1 · 摘自论文原文
  • 用冻结的视觉模型和少量样本提取动作时间线
  • 识别准确率与能力负相关,高能力者动作更复杂多样
  • 适合用于自动化评估护理培训中的真实表现

临床模拟培训中评估学习者能力依赖专家观察,耗时且易受主观影响。本文提出三阶段框架:(1)使用冻结的视觉编码器与少样本学习从第一人称护理模拟视频中提取动作时间线;(2)生成序列级特征与会话级识别指标;(3)关联这些指标与教师评分的能力等级。在22个密集标注会话(共3.8小时,493个动作)上,基于DINOv2的模型结合HMM Viterbi解码,在留一法1样本设置下达到57.4%的平均重叠率(MOF)。令人意外的是,识别准确率与能力呈负相关(mIoU的rho = -0.524,p = 0.012),该趋势在六种混杂因素控制下依然稳健:更熟练的学生表现出更多样、更难分类的动作流程。逐项分析显示,患者安全规程与团队沟通是最具代表性的行为模式。过程模型对比表明,高能力学生具有更一致的流程转换。结果提示,识别准确率可作为自动评估中补充性的教学信号。

原文摘要 · Abstract (English)

Assessing learner competency in clinical simulation requires expert observation that is time-intensive, difficult to scale, and subject to inter-rater variability. Vision-language models have emerged as a promising tool for understanding complex visual behavior. In this work, we investigate whether visual observations can provide educationally meaningful signals for competency assessment through a three-stage framework that (1) extracts action timelines from egocentric nursing simulation video using frozen visual encoders and few-shot learning, (2) derives sequence-level features and per-session recognition metrics, and (3) relates these to instructor-rated competency. Across 22 densely annotated sessions (3.8 hours, 493 actions), a frozen DINOv2 backbone with HMM Viterbi decoding achieves 57.4% MOF in leave-one-out 1-shot recognition. Surprisingly, we observe a negative trend between recognition accuracy and competency (rho = -0.524, p = 0.012 for mIoU), robust to six confound controls: more competent students produce diverse, harder-to-classify workflows, while simple sequence features show no such relationship. Per-item analysis identifies patient safety protocols and team communication as the expected behaviors most reflected in this pattern, and process model comparisons reveal that higher-competency students exhibit more protocol-consistent action transitions. These findings suggest that recognition accuracy may complement predicted action timelines as a pedagogically informative signal in automated competency assessment.

医学教育视觉分析能力评估视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。