arXiv:2605.12387cs.SDcs.LG2026-05

用半监督方法提升语音自信度检测,融合语义与声学特征。

A Semi-Supervised Framework for Speech Confidence Detection using Whisper

论文配图:A Semi-Supervised Framework for Speech Confidence Detection using Whisper
图 1 · 摘自论文原文
  • 结合Whisper语义嵌入与声学特征,构建混合检测框架。
  • 宏平均F1达0.751,少数类性能提升3%。
  • 高质量伪标签比大量数据增强更有效,适合小样本场景。

自动检测说话人自信度对自适应计算至关重要,但受限于标注数据稀缺和副语言标注的主观性。本文提出一种半监督混合框架,将Whisper编码器生成的深层语义嵌入与由eGeMAPS描述符及语音压力、不流畅性辅助概率估计构成的可解释声学特征向量融合。为减少对稀缺真实标签的依赖,引入不确定性感知的伪标签策略:模型为无标签数据生成标签,仅保留高质量样本用于训练。实验表明,该方法在宏平均F1上达到0.751,优于WavLM、HuBERT和Wav2Vec 2.0等自监督基线。混合架构也超越单一的Whisper基线,在少数类上实现3%的性能提升,证明显式韵律与辅助特征能提供深度语义表示中缺失的关键修正信号。消融实验进一步显示,经筛选的高置信度伪标签集优于盲目大规模增强,验证了数据质量优于数量对于感知自信度检测的重要性。

原文摘要 · Abstract (English)

Automatic detection of speaker confidence is critical for adaptive computing but remains constrained by limited labelled data and the subjectivity of paralinguistic annotations. This paper proposes a semi-supervised hybrid framework that fuses deep semantic embeddings from the Whisper encoder with an interpretable acoustic feature vector composed of eGeMAPS descriptors and auxiliary probability estimates of vocal stress and disfluency. To mitigate reliance on scarce ground truth data, we introduce an Uncertainty-Aware Pseudo-Labelling strategy where a model generates labels for unlabelled data, retaining only high-quality samples for training. Experimental results demonstrate that the proposed approach achieves a Macro-F1 score of 0.751, outperforming self-supervised baselines, including WavLM, HuBERT, and Wav2Vec 2.0. The hybrid architecture also surpasses the unimodal Whisper baseline, yielding a 3\% improvement in the minority class, confirming that explicit prosodic and auxiliary features provide necessary corrective signals which are otherwise lost in deep semantic representations. Ablation studies further show that a curated set of high confidence pseudo-labels outperforms indiscriminate large scale augmentation, confirming that data quality outweighs quantity for perceived confidence detection.

语音分析半监督学习自信度检测特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。