用视觉识别技术自动评估急救团队在混合现实训练中的操作表现。
Trainee Action Recognition through Interaction Analysis in CCATT Mixed-Reality Training
- 结合认知任务分析与多模态学习,构建可解释的评估框架。
- 通过改进的人体-物体交互模型识别设备操作,精准追踪反应时间与任务时长。
- 适合医疗训练评估、人机协同研究者参考。
本研究探讨了利用混合现实模拟器对重症航空转运团队(CCATT)成员进行训练的方法,该场景复现了空中医疗后送的高压环境。每个团队由医生、护士和呼吸治疗师组成,需在飞行过程中通过管理呼吸机、输液泵和吸引装置来稳定重伤士兵。熟练操作不仅依赖临床技能,还需具备情境意识、快速决策、有效沟通和协调任务管理等认知能力,且这些能力必须在压力下保持。近年来,仿真与多模态数据分析的进步使得更客观、全面的绩效评估成为可能;而传统导师评估主观性强,易遗漏关键事件,影响泛化性和一致性。然而,基于AI的自动化评估仍需人工标注以训练算法识别复杂团队动态,尤其在环境噪声干扰和多人追踪中需准确再识别。为此,本文提出一种系统性、数据驱动的评估框架,融合认知任务分析(CTA)与多模态学习分析(MMLA)。我们开发了适用于CCATT训练的领域特定CTA模型,并采用微调的人体-物体交互模型(级联解耦网络,CDN)构建基于视觉的动作识别流水线,实现对训练员-设备交互的时序检测与追踪。这些交互自动生成性能指标(如反应时间、任务持续时间),并映射至针对CCATT操作设计的层级化CTA模型,从而实现可解释、领域相关的绩效评估。
原文摘要 · Abstract (English)
This study examines how Critical Care Air Transport Team (CCATT) members are trained using mixed-reality simulations that replicate the high-pressure conditions of aeromedical evacuation. Each team - a physician, nurse, and respiratory therapist - must stabilize severely injured soldiers by managing ventilators, IV pumps, and suction devices during flight. Proficient performance requires clinical expertise and cognitive skills, such as situational awareness, rapid decision-making, effective communication, and coordinated task management, all of which must be maintained under stress. Recent advances in simulation and multimodal data analytics enable more objective and comprehensive performance evaluation. In contrast, traditional instructor-led assessments are subjective and may overlook critical events, thereby limiting generalizability and consistency. However, AI-based automated and more objective evaluation metrics still demand human input to train the AI algorithms to assess complex team dynamics in the presence of environmental noise and the need for accurate re-identification in multi-person tracking. To address these challenges, we introduce a systematic, data-driven assessment framework that combines Cognitive Task Analysis (CTA) with Multimodal Learning Analytics (MMLA). We have developed a domain-specific CTA model for CCATT training and a vision-based action recognition pipeline using a fine-tuned Human-Object Interaction model, the Cascade Disentangling Network (CDN), to detect and track trainee-equipment interactions over time. These interactions automatically yield performance indicators (e.g., reaction time, task duration), which are mapped onto a hierarchical CTA model tailored to CCATT operations, enabling interpretable, domain-relevant performance evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。