构建多模态临床手势数据集,助力手部动作质量评估
EHWGesture -- A dataset for multimodal understanding of clinical gestures
- 用双摄像头+事件相机采集1100+段多视角视频
- 提供精确手部关键点追踪与动作速度分级标注
- 适合临床手功能评估与多模态手势识别研究
手部动作理解对人机交互中的临床手部灵巧度自动评估至关重要。尽管深度学习提升了静态手势识别能力,动态手势因时空变化复杂仍具挑战。现有数据集普遍存在多模态多样性不足、多视角覆盖有限、精准轨迹标注缺失及动作质量标签缺乏等问题。本文提出EHWGesture,一个面向临床的多模态视频数据集,包含5种临床相关手势,共超过1100段记录(总时长约6小时),由25名健康受试者在双高分辨率RGB-Depth相机与事件相机下录制。通过运动捕捉系统获取精确的手部关键点真值轨迹,并实现所有设备的空间标定与同步,确保跨模态对齐。为嵌入动作质量评估任务,数据按执行速度分组,模拟临床手灵巧度评价标准。基线实验表明该数据集在手势分类、触发检测与动作质量评估方面具有潜力,可作为推进多模态临床手势理解的综合性基准。
原文摘要 · Abstract (English)
Hand gesture understanding is essential for several applications in human-computer interaction, including automatic clinical assessment of hand dexterity. While deep learning has advanced static gesture recognition, dynamic gesture understanding remains challenging due to complex spatiotemporal variations. Moreover, existing datasets often lack multimodal and multi-view diversity, precise ground-truth tracking, and an action quality component embedded within gestures. This paper introduces EHWGesture, a multimodal video dataset for gesture understanding featuring five clinically relevant gestures. It includes over 1,100 recordings (6 hours), captured from 25 healthy subjects using two high-resolution RGB-Depth cameras and an event camera. A motion capture system provides precise ground-truth hand landmark tracking, and all devices are spatially calibrated and synchronized to ensure cross-modal alignment. Moreover, to embed an action quality task within gesture understanding, collected recordings are organized in classes of execution speed that mirror clinical evaluations of hand dexterity. Baseline experiments highlight the dataset's potential for gesture classification, gesture trigger detection, and action quality assessment. Thus, EHWGesture can serve as a comprehensive benchmark for advancing multimodal clinical gesture understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。