arXiv:2501.08609cs.CV2025-01中稿 · publication in IEE…被引 2

用视频直接评估自闭症儿童动作模仿能力,无需人工标注和数据清洗。

Computerized Assessment of Motor Imitation for Distinguishing Autism in Video (CAMI-2DNet)

  • 基于编码器-解码器结构,从视频中提取与体型、视角无关的动作特征。
  • 在合成数据和真实数据上训练,相似度评分可区分自闭症儿童与正常儿童。
  • 比传统方法更自动化,适合大规模筛查,且效果接近三维动作捕捉方案。

动作模仿缺陷常出现在自闭症谱系障碍(ASC)个体中,可能作为表型用于缓解自闭症异质性。传统评估方法主观性强、耗时费力,需大量人工训练。现有计算机化评估方法如CAMI-3D(基于动作捕捉)和CAMI-2D(基于视频)虽较客观,但仍依赖繁琐的数据归一化、清洗及人工标注。为此,本文提出CAMI-2DNet,一种可扩展且可解释的深度学习方法,直接处理视频数据,无需数据预处理与标注。该模型采用编码器-解码器架构,将视频映射为与体形、视角等干扰因素解耦的动作编码。通过重排虚拟角色的动作、体形和摄像机视角生成合成数据,并结合真实参与者数据进行训练。通过计算个体间动作编码的相似度,实现对模仿能力的自动评估。对比分析表明,CAMI-2DNet与人工评分高度相关,且在区分ASC儿童与神经典型(NT)儿童方面优于CAMI-2D。其性能与CAMI-3D相当,但可直接处理视频,无需额外数据处理与人工标注,更具实用性。

原文摘要 · Abstract (English)

Motor imitation impairments are commonly reported in individuals with autism spectrum conditions (ASCs), suggesting that motor imitation could be used as a phenotype for addressing autism heterogeneity. Traditional methods for assessing motor imitation are subjective, labor-intensive, and require extensive human training. Modern Computerized Assessment of Motor Imitation (CAMI) methods, such as CAMI-3D for motion capture data and CAMI-2D for video data, are less subjective. However, they rely on labor-intensive data normalization and cleaning techniques, and human annotations for algorithm training. To address these challenges, we propose CAMI-2DNet, a scalable and interpretable deep learning-based approach to motor imitation assessment in video data, which eliminates the need for data normalization, cleaning and annotation. CAMI-2DNet uses an encoder-decoder architecture to map a video to a motion encoding that is disentangled from nuisance factors such as body shape and camera views. To learn a disentangled representation, we employ synthetic data generated by motion retargeting of virtual characters through the reshuffling of motion, body shape, and camera views, as well as real participant data. To automatically assess how well an individual imitates an actor, we compute a similarity score between their motion encodings, and use it to discriminate individuals with ASCs from neurotypical (NT) individuals. Our comparative analysis demonstrates that CAMI-2DNet has a strong correlation with human scores while outperforming CAMI-2D in discriminating ASC vs NT children. Moreover, CAMI-2DNet performs comparably to CAMI-3D while offering greater practicality by operating directly on video data and without the need for ad-hoc data normalization and human annotations.

自闭症检测动作模仿视频分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。