用数据驱动新方法解析面部表情,提升自闭症预测准确率
Beyond FACS: Data-driven Facial Expression Dictionaries, with Application to Predicting Autism
- 构建无监督的3D面部运动编码系统Facial Basis,替代传统FACS
- 在面对面和远程对话中预测自闭症表现优于主流AU检测器
- 支持完整面部动作重建,适合需全面分析表情的研究者
面部动作编码系统(FACS)被广泛用于研究面部行为与心理健康的关系。但其人工标注耗时费力,机器学习自动检测仍存在诸多限制:部分面部动作单元(AU)检测精度不足,且许多AU被排除在外,无法实现FACS对任意面部表情的完整表征。本文提出一种替代方案——数据驱动的‘Facial Basis’编码系统,该系统基于可解释的3D局部面部运动单元,克服了自动化FACS的三项结构性缺陷:第一,完全无监督,避免昂贵的人工标注;第二,能重构所有可观测的面部运动,而非仅限于特定可识别动作;第三,单元具有可加性,即使在非可加组合下仍能有效检测。实验表明,该方法在从面对面及远程对话中预测自闭症诊断方面,优于最常用的AU检测器。据我们所知,Facial Basis是首个可替代FACS、将视频中面部表情分解为局部运动的系统。代码已开源:github.com/sariyanidi/FacialBasis。
原文摘要 · Abstract (English)
The Facial Action Coding System (FACS) has been used by numerous studies to investigate the links between facial behavior and mental health. The laborious and costly process of FACS coding has motivated the development of machine learning frameworks for Action Unit (AU) detection. Despite intense efforts spanning three decades, the detection accuracy for many AUs is considered to be below the threshold needed for behavioral research. Also, many AUs are excluded altogether, making it impossible to fulfill the ultimate goal of FACS-the representation of any facial expression in its entirety. This paper considers an alternative approach. Instead of creating automated tools that mimic FACS experts, we propose to use a new coding system that mimics the key properties of FACS. Specifically, we construct a data-driven coding system called the Facial Basis, which contains units that correspond to localized and interpretable 3D facial movements, and overcomes three structural limitations of automated FACS coding. First, the proposed method is completely unsupervised, bypassing costly, laborious and variable manual annotation. Second, Facial Basis reconstructs all observable movement, rather than relying on a limited repertoire of recognizable movements (as in automated FACS). Finally, the Facial Basis units are additive, whereas AUs may fail detection when they appear in a non-additive combination. The proposed method outperforms the most frequently used AU detector in predicting autism diagnosis from in-person and remote conversations, highlighting the importance of encoding facial behavior comprehensively. To our knowledge, Facial Basis is the first alternative to FACS for deconstructing facial expressions in videos into localized movements. We provide an open source implementation of the method at github.com/sariyanidi/FacialBasis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。