通过分析自拍视频中的微表情动态,无语言地筛查早期痴呆症。
Passive Dementia Screening via Facial Temporal Micro-Dynamics Analysis of In-the-Wild Talking-Head Video
- 提取面部细微动作时间序列,以运动分布而非强度判断异常。
- 在公开数据集上实现0.953的AUROC和0.851的F1-score。
- 无需语音或脚本,适用于真实场景的跨设备、跨文化筛查。
本文针对从短时摄像头自拍说话视频中被动筛查痴呆症,提出一种无需语言的面部时间微动态分析方法,用于检测早期神经认知变化。该方法可在自然状态下大规模分析未经剪辑的视频,捕捉真实面部行为,且不依赖临床干预,具备跨设备、跨话题与跨文化的可迁移性。现有资源多聚焦语音或剧本化访谈,限制了其在非临床环境的应用,并将预测结果与语言内容绑定。相比之下,本文研究仅依靠面部时间动态(包括眨眼模式、微小嘴部下颌运动、凝视变动性及细微头部调整)是否足以实现痴呆筛查。通过稳定面部信号,将微动作转化为可解释的时间序列,进行平滑处理,并将短时段信息汇总为片段级统计量。每个窗口由其运动成分占比(各动作流相对贡献)编码,使模型关注运动分布而非幅度,增强可解释性。同时引入新数据集YT DemTalk,包含300段公开视频(150例自报痴呆者,150例对照),提供首个基准测试。在该数据集上,消融实验表明凝视不稳定性与口/下颌动态最具判别力,轻量浅层分类器即可达到AUROC 0.953、平均精度AP 0.961、F1分数0.851、准确率0.857。
原文摘要 · Abstract (English)
We target passive dementia screening from short camera-facing talking head video, developing a facial temporal micro dynamics analysis for language free detection of early neuro cognitive change. This enables unscripted, in the wild video analysis at scale to capture natural facial behaviors, transferrable across devices, topics, and cultures without active intervention by clinicians or researchers during recording. Most existing resources prioritize speech or scripted interviews, limiting use outside clinics and coupling predictions to language and transcription. In contrast, we identify and analyze whether temporal facial kinematics, including blink dynamics, small mouth jaw motions, gaze variability, and subtle head adjustments, are sufficient for dementia screening without speech or text. By stabilizing facial signals, we convert these micro movements into interpretable facial microdynamic time series, smooth them, and summarize short windows into compact clip level statistics for screening. Each window is encoded by its activity mix (the relative share of motion across streams), thus the predictor analyzes the distribution of motion across streams rather than its magnitude, making per channel effects transparent. We also introduce YT DemTalk, a new dataset curated from publicly available, in the wild camera facing videos. It contains 300 clips (150 with self reported dementia, 150 controls) to test our model and offer a first benchmarking of the corpus. On YT DemTalk, ablations identify gaze lability and mouth/jaw dynamics as the most informative cues, and light weighted shallow classifiers could attain a dementia prediction performance of (AUROC) 0.953, 0.961 Average Precision (AP), 0.851 F1-score, and 0.857 accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。