用人脸和语言信息联合诊断睡眠呼吸暂停,准确率超91%。
An Attentive Dual-Encoder Framework Leveraging Multimodal Visual and Semantic Information for Automatic OSAHS Diagnosis

- 双编码器融合视觉与文本特征,注意力网格聚焦关键面部区域。
- 四分类严重程度识别达91.3%准确率,优于现有方法。
- 适合医学影像分析、多模态学习研究者参考。
阻塞性睡眠呼吸暂停低通气综合征(OSAHS)是因上气道阻塞导致缺氧和睡眠中断的常见睡眠障碍。传统多导睡眠图(PSG)诊断成本高、耗时长且体验差。现有基于面部图像的深度学习方法因特征捕捉不足和样本量有限,准确性受限。为此,我们提出一种融合视觉与语言输入的多模态双编码器模型。通过randomOverSampler平衡数据,利用注意力网格提取关键面部特征,并将生理数据转化为语义文本。跨注意力机制融合图像与文本信息以增强特征表达,有序回归损失保障训练稳定性。该方法在四分类严重程度任务中达到91.3%的顶1准确率,性能达当前最优。代码将在录用后公开。
原文摘要 · Abstract (English)
Obstructive sleep apnea-hypopnea syndrome (OSAHS) is a common sleep disorder caused by upper airway blockage, leading to oxygen deprivation and disrupted sleep. Traditional diagnosis using polysomnography (PSG) is expensive, time-consuming, and uncomfortable. Existing deep learning methods using facial image analysis lack accuracy due to poor facial feature capture and limited sample sizes. To address this, we propose a multimodal dual encoder model that integrates visual and language inputs for automated OSAHS diagnosis. The model balances data using randomOverSampler, extracts key facial features with attention grids, and converts physiological data into meaningful text. Cross-attention combines image and text data for better feature extraction, and ordered regression loss ensures stable learning. Our approach improves diagnostic efficiency and accuracy, achieving 91.3% top-1 accuracy in a four-class severity classification task, demonstrating state-of-the-art performance. Code will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。