arXiv:2604.22036cs.CVcs.AI2026-04

EgoMAGIC提供50种医疗操作的头戴式视频数据,用于训练医学视觉算法。

EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms

论文配图:EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms
图 1 · 摘自论文原文
  • 采集50项医疗任务的头戴摄像头视频,每项至少50条带标注数据。
  • 用195万标注训练40个YOLO模型,124类医疗物品检测平均mAP达0.526。
  • 适合研究医疗动作识别、错误检测与增强现实辅助系统开发。

本文介绍EgoMAGIC(Medical Assistance, Guidance, Instruction, and Correction)数据集,该数据集是美国国防高级研究计划局(DARPA)感知任务引导(PTG)项目的一部分。数据集包含50项医疗任务的3,355段视频,每项任务至少50条标注视频。这些视频主要通过集成音频的头戴式双目相机录制,旨在支持增强现实眼镜中虚拟助手的研发。为推动研究,数据集已公开,并配套开展针对8项医疗任务的动作检测挑战赛。基于该数据集,使用195万标注训练了40个YOLO模型,实现对124种医疗物品的检测。本文还提供了三种模型在8项任务上的基准结果,最优方法平均mAP达到0.526。尽管以动作检测为主要评测任务,该数据集同样适用于动作识别、物体识别与检测、错误检测等计算机视觉任务。数据集可通过zenodo.org(DOI: 10.5281/zenodo.19239154)获取。

原文摘要 · Abstract (English)

This paper introduces EgoMAGIC (Medical Assistance, Guidance, Instruction, and Correction), an egocentric medical activity dataset collected as part of DARPA's Perceptually-enabled Task Guidance (PTG) program. This dataset comprises 3,355 videos of 50 medical tasks, with at least 50 labeled videos per task. The primary objective of the PTG program was to develop virtual assistants integrated into augmented reality headsets to assist users in performing complex tasks. To encourage exploration and research using this dataset, the medical training data has been released along with an action detection challenge focused on eight medical tasks. The majority of the videos were recorded using a head-mounted stereo camera with integrated audio. From this dataset, 40 YOLO models were trained using 1.95 million labels to detect 124 medical objects, providing a robust starting point for developers working on medical AI applications. In addition to introducing the dataset, this paper presents baseline results on action detection for the eight selected medical tasks across three models, with the best-performing method achieving average mAP 0.526. Although this paper primarily addresses action detection as the benchmark, the EgoMAGIC dataset is equally suitable for action recognition, object identification and detection, error detection, and other challenging computer vision tasks. The dataset is accessible via zenodo.org (DOI: 10.5281/zenodo.19239154).

医疗视觉动作检测头戴数据计算机视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。