arXiv:2606.25606cs.CVcs.AI2026-06中稿 · the 2023 11th Inte…被引 4

用视频分析抑郁症严重程度,让模型决策可解释。

Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis

论文配图:Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis
图 1 · 摘自论文原文
  • 基于面部视频微调预训练模型,分析表情时序特征
  • 在AVEC数据集上提升预测性能,优于单脸基准
  • 生成视觉与量化解释,助力临床医生理解模型判断

抑郁症广泛流行,亟需客观、可解释的早期诊断手段。现有自动诊断方法虽有进步,但仍缺乏情感特异性与可解释性,尤其在视频中通过深度模型分析时序行为的解释能力不足。本文提出一种新框架,对在面部视频上训练的深度神经网络进行解释,聚焦于抑郁症严重程度的自动诊断。通过在AVEC抑郁症数据集的面部视频上微调预训练于动作识别的数据集的深度卷积神经网络(DCNN),该框架能分析面部区域与时序表情语义的显著性图。所提方法生成可视化与量化解释,揭示模型决策逻辑。同时,该视频建模方法超越了以往单脸基准,在预测性能上取得提升。整体表明,该框架不仅能从模型决策中生成可解释假设,还能增强抑郁症预测能力。

原文摘要 · Abstract (English)

Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for early diagnosis and intervention. As a multidisciplinary topic, these objective measures should be interpretable and accessible to health care professionals, ensuring effective collaboration and treatment planning in the realm of mental health care. Even though current automated depression diagnosis approaches improved over the last decade, a critical gap exists as they often lack affect-specificity and interpretability, limiting their practical application and potential impact on mental health care. In particular, interpretability from temporal activities from videos when deep models are used is not fully explored. In this study, we present a novel framework for analyzing Deep Neural Networks' decisions when trained on facial videos, specifically focusing on automatic depression severity diagnosis. By fine-tuning Deep Convolutional Neural Networks (DCNN) pre-trained on Action Recognition datasets on depression severity facial videos from AVEC depression dataset, our framework is able to interpret the model's saliency maps by examining face regions and temporal expression semantics. Our approach generates both visual and quantitative explanations for the model's decisions, providing greater insight into its reasoning. In addition to this interpretability, our video-based modeling has improved upon previous single-face benchmarks for visual depression diagnosis, resulting in enhanced predictive performance. Overall, our work demonstrates the successful development of a framework capable of generating hypotheses from a facial model's decisions while simultaneously improving depression's predictive capabilities.

抑郁症诊断视频分析可解释AI深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。