arXiv:2410.09289cs.SDcs.AI2024-10被引 3

用分层融合的Transformer模型,提升语音疾病预测准确率。

Multimodal Audio-based Disease Prediction with Transformer-based Hierarchical Fusion Network

  • 分层融合音频模态内与模态间特征,捕捉互补信息
  • 在新冠、帕金森和构音障碍预测上达顶尖水平
  • 适合医疗语音分析与多模态机器学习研究者

基于语音的疾病预测正成为传统医疗诊断的有力补充,有助于实现早期、便捷、无创的疾病检测与预防。多模态融合通过整合不同生物声学模态或域内的特征,已被证明能有效提升诊断性能。然而,现有方法多采用单一融合策略,仅关注模态内或模态间融合,限制了对多样化声学特征域及生物声学模态间互补性的充分挖掘。此外,对模态特异与模态共享空间中潜在依赖关系的探索不足,制约了其处理多模态数据固有异质性能力。为此,我们提出一种面向通用多模态语音疾病预测的基于Transformer的分层融合网络。该模型以层次化方式无缝集成模态内与模态间融合,并分别高效编码必要的模态内与模态间互补相关性。大量实验表明,该模型在预测三种疾病——新冠、帕金森病和病理型构音障碍——上达到当前最优性能,展现出在广泛语音疾病预测任务中的巨大潜力。此外,详尽的消融实验与定性分析验证了模型各核心组件的重要贡献。

原文摘要 · Abstract (English)

Audio-based disease prediction is emerging as a promising supplement to traditional medical diagnosis methods, facilitating early, convenient, and non-invasive disease detection and prevention. Multimodal fusion, which integrates features from various domains within or across bio-acoustic modalities, has proven effective in enhancing diagnostic performance. However, most existing methods in the field employ unilateral fusion strategies that focus solely on either intra-modal or inter-modal fusion. This approach limits the full exploitation of the complementary nature of diverse acoustic feature domains and bio-acoustic modalities. Additionally, the inadequate and isolated exploration of latent dependencies within modality-specific and modality-shared spaces curtails their capacity to manage the inherent heterogeneity in multimodal data. To fill these gaps, we propose a transformer-based hierarchical fusion network designed for general multimodal audio-based disease prediction. Specifically, we seamlessly integrate intra-modal and inter-modal fusion in a hierarchical manner and proficiently encode the necessary intra-modal and inter-modal complementary correlations, respectively. Comprehensive experiments demonstrate that our model achieves state-of-the-art performance in predicting three diseases: COVID-19, Parkinson's disease, and pathological dysarthria, showcasing its promising potential in a broad context of audio-based disease prediction tasks. Additionally, extensive ablation studies and qualitative analyses highlight the significant benefits of each main component within our model.

语音诊断多模态融合Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。