用固定摄像头实现可扩展的护理模拟实时分析,提升行为识别与场景理解能力。
Toward Scalable Co-located Practical Learning: Assisting with Computer Vision and Multimodal Analytics
- 通过两阶段迁移学习适应摄像头微调,提升目标检测精度
- 2022年数据集上mAP50达0.850,性能优于原始模型
- 支持按行为-区域关联进行可检索的复盘,适合教学反馈
共处式实践学习在患者周围留下可见动作、任务资源和房间区域的痕迹,但这些痕迹常需现场观察或事后视频回溯。固定广角摄像头可减轻感知负担,但回溯流程不仅需检测行为,还需在摄像头轻微位移后保持检测稳定性,将检测到的行为轨迹与教师标注结果关联,并保留房间区域上下文。本研究在重复护理模拟中评估了固定摄像头流程。采用统一六码分类体系,测试了YOLO26仅目标训练及两阶段源到目标适配,在两个同室侧视数据源上进行验证。将51个教师标注会话的检测结果转化为一秒级行为与行为-区域轨迹,用于率、有序网络、转移网络及序列分析。两阶段适配使2021年目标视图的平均mAP50从0.815提升至0.848,2022年目标视图从0.690升至0.855;在平衡目标样本量N=22时,2022年模型达到0.850 mAP50。行为轨迹分析显示,低任务表现会话中手机使用更高。区域标签改变了对患者互动的解读:高绩效会话中主要护理区互动更强,低绩效会话中次级区域互动更突出。有序与转移网络模型表明,区域顺序关系贡献超越行为频率,最强任务绩效分类器结合了区域与共现特征。最终生成的轨迹最适用于可搜索的模拟复盘,供教师审查检测时刻而非接收自动化评分。
原文摘要 · Abstract (English)
Co-located practical learning leaves evidence in visible actions around patients, task resources and room zones, but these traces are often recovered through live observation or retrospective video review. Fixed wide-angle video could reduce sensing burden, yet a debriefing pipeline must do more than detect behaviours: it must maintain detection after small camera-position shifts, relate the detector-derived behaviour trace to instructor-labelled outcomes and preserve room-zone context. This study evaluates a fixed-camera pipeline in repeated nursing simulation. Using a harmonised six-code taxonomy, we tested YOLO26 target-only training and two-stage source-to-target adaptation across two same-room side-view data sources. We then converted detections from 51 instructor-labelled sessions into one-second behaviour and behaviour-zone traces for rate, ordered-network, transition-network and sequence analyses. Two-stage adaptation improved mean mAP50 from 0.815 to 0.848 for the 2021 target view and from 0.690 to 0.855 for the smaller 2022 target view; with a balanced target quota of \(N = 22\), the 2022 model reached 0.850 mAP50. In the detector-derived behaviour trace analyses, higher phone use characterised low task-performance sessions. Zone labels changed the interpretation of patient interaction: primary patient-care-zone interaction was stronger in higher-performance sessions, while secondary-zone interaction was stronger in lower-performance sessions. Ordered and transition network models showed that ordered room-zone relations contributed beyond behaviour frequency, with the strongest task-performance classifier using zoned and co-presence features. The resulting trace is most appropriate for searchable simulation debriefing, where instructors inspect detected moments rather than receive automated assessment scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。