arXiv:2607.08714eess.AS2026-07

用语音+临床数据融合检测哮喘,手机录音即可筛查。

Multimodal Digital Biomarker for Asthma: Complementary Roles of Vocal, Clinical and Demographic Factors

论文配图:Multimodal Digital Biomarker for Asthma: Complementary Roles of Vocal, Clinical and Demographic Factors
图 1 · 摘自论文原文
  • 多模态专家混合模型融合语音与临床信息
  • 准确率AUROC达0.85,优于单一模式
  • 能根据症状轻重自动调整依赖特征,适合基层推广

哮喘全球影响超2.6亿人,但诊断仍依赖肺功能测试和专科评估,限制了初级医疗和资源匮乏地区的可及性。语音生物标志物提供了一种无创替代方案,但以往研究多仅关注声学特征而未整合临床背景。本文提出一种多模态混合专家框架,自适应融合持续元音发音和朗读文本任务的声学嵌入,以及结构化临床与人口统计学数据。模型在1,218例哮喘患者与健康对照组成的Colive Voice队列上进行评估,多模态模型达到AUROC 0.85、Brier分数0.17,优于单模态和双模态方法。自适应门控分析显示,呼吸症状负担较重者更依赖语音特征,而症状较轻者临床特征贡献更大。这些发现支持利用智能手机采集的语音实现可扩展且可解释的哮喘筛查。

原文摘要 · Abstract (English)

Asthma affects over 260 million people worldwide, yet diagnosis remains dependent on spirometry and specialist assessment, limiting accessibility in primary care and low-resource settings. Vocal biomarkers offer a promising non-invasive alternative, but prior studies have largely focused on acoustic features without integrating clinical context. We present a multimodal Mixture-of-Experts framework for asthma detection that adaptively combines acoustic embeddings from sustained vowel phonation and reading passage tasks with structured clinical and demographic data. The model was evaluated on a matched cohort of 1,218 asthma cases and healthy controls from the Colive Voice study. The multimodal model achieved an AUROC of 0.85 and Brier score of 0.17, outperforming unimodal and bimodal approaches. Adaptive gating analysis revealed increased reliance on audio features in participants with greater respiratory symptom burden, whereas clinical features contributed more strongly in less symptomatic individuals. These findings support scalable and explainable asthma screening using smartphone-collected voice recordings.

哮喘筛查语音生物标志物多模态学习智能手机应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。