融合视觉与结构化特征,提升面瘫检测准确率
A Multimodal Fusion Model Leveraging MLP Mixer and Handcrafted Features-based Deep Learning Networks for Facial Palsy Detection
- 用MLP混洗器处理图像,全连接网络处理面部坐标等结构数据
- 多模态融合模型F1达96.00,优于单一模态模型
- 适合医疗AI开发者和临床辅助诊断研究者
算法化面瘫检测有望改善当前依赖临床医生主观、耗时评估的现状。本文提出一种基于多模态融合的深度学习模型:使用MLP混洗器处理未结构化数据(如RGB图像或含面部轮廓线的图像),同时采用前馈神经网络处理结构化数据(如面部关键点坐标、表情特征或人工设计特征),以实现面瘫检测。我们通过分析20名面瘫患者和20名健康受试者的视频数据,研究不同数据模态的影响及多模态融合的优势。结果表明,该多模态融合模型达到96.00的F1分数,显著高于仅使用人工特征训练的前馈神经网络(82.80 F1)和仅使用原始RGB图像训练的MLP混洗器模型(89.00 F1)。
原文摘要 · Abstract (English)
Algorithmic detection of facial palsy offers the potential to improve current practices, which usually involve labor-intensive and subjective assessments by clinicians. In this paper, we present a multimodal fusion-based deep learning model that utilizes an MLP mixer-based model to process unstructured data (i.e. RGB images or images with facial line segments) and a feed-forward neural network to process structured data (i.e. facial landmark coordinates, features of facial expressions, or handcrafted features) for detecting facial palsy. We then contribute to a study to analyze the effect of different data modalities and the benefits of a multimodal fusion-based approach using videos of 20 facial palsy patients and 20 healthy subjects. Our multimodal fusion model achieved 96.00 F1, which is significantly higher than the feed-forward neural network trained on handcrafted features alone (82.80 F1) and an MLP mixer-based model trained on raw RGB images (89.00 F1).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。