arXiv:2505.20050eess.AScs.CL2025-05中稿 · Interspeech 2025被引 2

用Transformer融合多源语音数据,提升嗓音疾病检测准确率

MVP: Multi-source Voice Pathology detection

  • 直接处理原始语音信号,用Transformer融合朗读与长音录制数据
  • 中间特征融合使不同语音类型互补,最高提升13%的AUC
  • 适合医疗语音分析、跨语言嗓音病筛查的研究者

嗓音障碍严重影响患者生活质量,但非侵入式自动化诊断因病理语音数据稀缺和录音来源多样而进展缓慢。本文提出MVP(多源嗓音疾病检测),一种直接在原始语音信号上运行的Transformer方法。通过波形拼接、中间特征融合与决策级组合三种策略,融合朗读句子与持续元音录音。在德语、葡萄牙语和意大利语数据集上的实验证明,中间特征融合能最好地捕捉两类录音的互补特性,相比单源方法最高提升13% AUC。

原文摘要 · Abstract (English)

Voice disorders significantly impact patient quality of life, yet non-invasive automated diagnosis remains under-explored due to both the scarcity of pathological voice data, and the variability in recording sources. This work introduces MVP (Multi-source Voice Pathology detection), a novel approach that leverages transformers operating directly on raw voice signals. We explore three fusion strategies to combine sentence reading and sustained vowel recordings: waveform concatenation, intermediate feature fusion, and decision-level combination. Empirical validation across the German, Portuguese, and Italian languages shows that intermediate feature fusion using transformers best captures the complementary characteristics of both recording types. Our approach achieves up to +13% AUC improvement over single-source methods.

嗓音检测多源融合Transformer医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。