通过解耦多模态特征,提升神经疾病评估的准确与可解释性。
DIVINE: Coordinating Multimodal Disentangled Representations for Oro-Facial Neurological Disorder Assessment
- 解耦共享与模态特异性表示,增强模型可解释性。
- 在多模态和单模态条件下均达98.26%准确率,优于基线。
- 适合临床辅助诊断与跨模态医学分析研究者使用。
本研究提出一种多模态框架,通过捕捉语音与面部线索来预测口面神经障碍。我们假设,在多模态基础模型嵌入中显式解耦共享与模态特异性表示,可提升临床可解释性与泛化能力。为此,我们提出DIVINE——一个完全解耦的多模态框架,基于SOTA音频与视频基础模型提取的表示,结合分层变分瓶颈、稀疏门控融合与可学习症状标记。该框架采用多任务学习,联合预测诊断类别(健康对照、肌萎缩侧索硬化症、中风)与严重程度(轻度、中度、重度)。模型使用同步音视频输入训练,并在多伦多神经面数据集上,于全模态(音视频)及单模态(仅音频、仅视频)测试条件下评估。结果表明,采用DeepSeek-VL2与TRILLsson组合的DIVINE达到98.26%准确率与97.51% F1分数,为当前最优。在模态受限场景下仍表现优异,展现出强泛化能力,显著优于单模态模型与基线融合方法。据我们所知,DIVINE是首个结合跨模态解耦、自适应融合与多任务学习的神经障碍综合评估框架。
原文摘要 · Abstract (English)
In this study, we present a multimodal framework for predicting neuro-facial disorders by capturing both vocal and facial cues. We hypothesize that explicitly disentangling shared and modality-specific representations within multimodal foundation model embeddings can enhance clinical interpretability and generalization. To validate this hypothesis, we propose DIVINE a fully disentangled multimodal framework that operates on representations extracted from state-of-the-art (SOTA) audio and video foundation models, incorporating hierarchical variational bottlenecks, sparse gated fusion, and learnable symptom tokens. DIVINE operates in a multitask learning setup to jointly predict diagnostic categories (Healthy Control,ALS, Stroke) and severity levels (Mild, Moderate, Severe). The model is trained using synchronized audio and video inputs and evaluated on the Toronto NeuroFace dataset under full (audio-video) as well as single-modality (audio-only and video-only) test conditions. Our proposed approach, DIVINE achieves SOTA result, with the DeepSeek-VL2 and TRILLsson combination reaching 98.26% accuracy and 97.51% F1-score. Under modality-constrained scenarios, the framework performs well, showing strong generalization when tested with video-only or audio-only inputs. It consistently yields superior performance compared to unimodal models and baseline fusion techniques. To the best of our knowledge, DIVINE is the first framework that combines cross-modal disentanglement, adaptive fusion, and multitask learning to comprehensively assess neurological disorders using synchronized speech and facial video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。