用多层级语音建模提升罕见病语音检测,仅需11个标注样本就达90%性能。
Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling
- 从帧、片段到会话层联合建模语音特征,动态融合多粒度信息。
- 在仅11个标注样本下达到全监督90%的准确率,数据效率极高。
- 适用于跨语言、跨病症的医疗语音分析,适合资源稀缺场景。
从语音中检测医学状况本质上是弱监督学习问题:单一且常有噪声的会话级标签需对应长而复杂的音频中的细微模式。该任务因数据极度匮乏及临床标注的主观性而更加困难。虽然半监督学习(SSL)能利用未标注数据,但现有音频方法未能解决病理特征在患者语音中非均匀分布的核心挑战。本文提出一种全新的纯音频半监督框架,通过在未分割的临床对话中联合学习帧级、段级和会话级表示,显式建模这一层次结构。端到端方法动态聚合多粒度特征,并生成高质量伪标签以高效利用未标注数据。大量实验表明,该框架模型无关、跨语言与病症鲁棒,且高度数据高效——例如,仅用11个标注样本即达到全监督90%的性能。本工作为医疗语音分析中弱远端监督的学习提供了原则性解决方案。代码已开源:https://github.com/fispresent/semi_pathological。
原文摘要 · Abstract (English)
Detecting medical conditions from speech acoustics is fundamentally a weakly-supervised learning problem: a single, often noisy, session-level label must be linked to nuanced patterns within a long, complex audio recording. This task is further hampered by severe data scarcity and the subjective nature of clinical annotations. While semi-supervised learning (SSL) offers a viable path to leverage unlabeled data, existing audio methods often fail to address the core challenge that pathological traits are not uniformly expressed in a patient's speech. We propose a novel, audio-only SSL framework that explicitly models this hierarchy by jointly learning from frame-level, segment-level, and session-level representations within unsegmented clinical dialogues. Our end-to-end approach dynamically aggregates these multi-granularity features and generates high-quality pseudo-labels to efficiently utilize unlabeled data. Extensive experiments show the framework is model-agnostic, robust across languages and conditions, and highly data-efficient-achieving, for instance, 90% of fully-supervised performance using only 11 labeled samples. This work provides a principled approach to learning from weak, far-end supervision in medical speech analysis. The code is available at https://github.com/fispresent/semi_pathological.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。