用视频识别新生儿护理操作,提升医护记录效率
Infant Care Video Dataset for Classification of Interventions Using Transformers

- 构建4144段模拟婴儿护理视频数据集,涵盖12类干预动作
- 基于Transformer模型实现93.97%准确率,显著优于逐帧分析
- 支持多角度与肤色条件测试,适合医疗自动化系统研发
新生儿重症监护室(NICU)的医疗记录存在重大挑战,护士约25%的工作时间用于文书工作,高达60%的干预措施未被记录。为实现从视频中自动检测干预行为,我们提出婴儿护理视频数据集(ICVD),包含4,144段视频,覆盖12种模拟干预类别,旨在开发自动化文档系统。采用人形模拟器的方法,系统性地改变摄像头角度和医护人员肤色等条件,同时确保隐私合规。利用视频变换器架构(TimeSformer和MotionFormer),在12类婴儿护理任务上达到93.97%和93.17%的顶级分类准确率。消融实验表明,与逐帧方法(23.17%准确率)相比,使用时序建模可提升70.80%性能,验证了时序建模的必要性。ICVD为降低新生儿护理环境中的临床负担、改进现有实践提供了坚实基础。
原文摘要 · Abstract (English)
Healthcare documentation in the neonatal intensive care unit (NICU) presents significant challenges, with nurses spending approximately 25\% of their time on record-keeping, while up to 60\% of interventions remain undocumented. Motivated by the need to detect interventions from video automatically, we present the Infant Care Video Dataset (ICVD), a collection of 4,144 videos spanning 12 simulated intervention classes designed for developing automated documentation systems. Our manikin-based approach systematically varies conditions, such as camera angle and clinician skin tone, while ensuring privacy compliance. Using video transformer architectures (TimeSformer and MotionFormer), we establish strong baseline performance (93.97\% and 93.17\% top-1 accuracy) among the 12 infant care classes. Our ablation study comparing temporal models with a framewise approach (23.17\% accuracy) demonstrates a 70.80\% performance gap, validating the need for temporal modeling. The ICVD provides a foundation for developing automated documentation systems to reduce clinical burden in neonatal care environments and improve existing practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。