arXiv:2606.16505cs.SDcs.LG2026-06被引 1

用伪标签和Whisper嵌入提升语音自信度检测准确率

Semi-Supervised Speech Confidence Detection using Pseudo-Labelling and Whisper Embeddings

论文配图:Semi-Supervised Speech Confidence Detection using Pseudo-Labelling and Whisper Embeddings
图 1 · 摘自论文原文
  • 结合人工特征与Whisper嵌入,通过共注意力融合多模态信息
  • 在有限标注数据下实现75%的识别准确率
  • 适合教育场景中个性化口语反馈系统开发

理解说话人自信程度在教育环境中至关重要,有助于提供个性化反馈并提升学习效果。本文提出一种新框架,通过整合人工设计的语音特征(如音高、音量、语速、不流畅性和压力)与Whisper编码器生成的嵌入表示,实现说话人自信度检测。为缓解标注数据不足问题,采用伪标签技术扩充标注数据集,使模型能同时学习人工标注和模型生成的标签。框架利用共注意力机制融合不同来源的特征表示,最终达到75%的整体准确率。该研究推动了语音分析的发展,支持个性化学习与口语能力培养的应用。

原文摘要 · Abstract (English)

Understanding speaker confidence is crucial in educational settings, as it can enhance personalised feedback and improve learning outcomes. This study introduces a novel framework for detecting speaker confidence by integrating human-engineered features with embeddings from the Whisper encoder. To address data limitations, a pseudo-labelling technique is employed to expand the labelled dataset, allowing the model to learn from both human-annotated and model-generated labels. The framework combines traditional speech features including pitch, volume, rate of speech, and the presence of disfluencies and stress, with Whisper embeddings, and uses a co-attention mechanism to fuse these representations and achieve an overall accuracy of 75%. This study contributes to advancing speech analysis, enabling applications that support personalised learning and speaking skill development.

语音分析自信检测伪标签Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。