arXiv:2601.22319eess.ASeess.SP2026-01中稿 · IEEE ICASSP 2026被引 1

优化语音医疗分类的自监督学习,提升小数据下的疾病识别效果。

Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification

  • 采用内容感知掩码和局部归一化,增强病理语音特征捕捉。
  • 使用平均绝对误差损失,使模型在小样本下更稳定,F1达0.688。
  • 适合医疗语音分析、小样本临床建模的研究者参考。

人类语音是一种有前景的非侵入式数字生物标志物,但基于语音的健康分析受制于数据稀缺和领域不匹配问题——在通用音频上预训练的模型难以捕捉临床语音中的细微病理特征。本文研究了基于掩码自编码器(MAE)的领域自适应自监督学习,发现标准配置对医疗音频效果不佳。利用多机构的病理语音数据集Bridge2AI-Voice,系统考察了三个关键因素:重建损失(均方误差与平均绝对误差)、归一化方式(块级与全局)、掩码策略(随机与内容感知)。最优设计结合平均绝对误差损失、块级归一化与内容感知掩码,在10次微调中取得0.688±0.009的宏平均F1,优于在大规模通用音频上预训练的强基线(0.663±0.011)。结果表明,平均绝对误差损失提升鲁棒性,内容感知掩码通过强调信息丰富区域显著提升性能。这凸显了在数据受限的医疗音频应用中组件级优化的重要性。

原文摘要 · Abstract (English)

The human voice is a promising non-invasive digital biomarker, yet deep learning for voice-based health analysis is hindered by data scarcity and domain mismatch, where models pre-trained on general audio fail to capture the subtle pathological features characteristic of clinical voice data. To address these challenges, we investigate domain-adaptive self-supervised learning (SSL) with Masked Autoencoders (MAE) and demonstrate that standard configurations are suboptimal for health-related audio. Using the Bridge2AI-Voice dataset, a multi-institutional collection of pathological voices, we systematically examine three performance-critical factors: reconstruction loss (Mean Absolute Error vs. Mean Squared Error), normalization (patch-wise vs. global), and masking (random vs. content-aware). Our optimized design, which combines Mean Absolute Error (MA-Error) loss, patch-wise normalization, and content-aware masking, achieves a Macro F1 of $0.688 \pm 0.009$ (over 10 fine-tuning runs), outperforming a strong out-of-domain SSL baseline pre-trained on large-scale general audio, which has a Macro F1 of $0.663 \pm 0.011$. The results show that MA-Error loss improves robustness and content-aware masking boosts performance by emphasizing information-rich regions. These findings highlight the importance of component-level optimization in data-constrained medical applications that rely on audio data.

语音医疗自监督学习小样本医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。