arXiv:2409.16937eess.AScs.AI2024-09中稿 · ICASSP 2025被引 6

用声音和语言双重特征自动生成高置信度标签,少用标注数据也能精准识别认知状态。

Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling

  • 融合声学与语言特征,通过双模态一致性筛选高置信度伪标签。
  • 仅用30%标注数据,情绪识别与痴呆检测性能接近全监督模型。
  • 适合标注成本高、需主观评估的语音认知分析任务。

语音分类任务中缺乏标注数据是普遍难题,尤其在需要大量主观评估的认知状态识别任务中。本文提出一种半监督学习框架,引入新颖的多视图伪标签方法,同时利用声学与语言特征选择高置信度数据用于训练。声学方面,通过多个音频编码器生成嵌入,计算未标注数据与已标注数据之间的弗雷歇音频距离;语言方面,使用大语言模型对自动语音识别转录进行修正,并基于特定任务知识预测标签。当两种来源的伪标签一致时,判定为高置信度数据,不一致则视为低置信度。随后使用双模态分类器迭代标注低置信度数据,直至满足预设条件。我们在情绪识别与痴呆检测任务上评估该框架,实验表明,仅使用30%标注数据即可达到与全监督学习相当的性能,显著优于两个基线方法。

原文摘要 · Abstract (English)

The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling method that leverages both acoustic and linguistic characteristics to select the most confident data for training the classification model. Acoustically, unlabeled data are compared to labeled data using the Frechet audio distance, calculated from embeddings generated by multiple audio encoders. Linguistically, large language models are prompted to revise automatic speech recognition transcriptions and predict labels based on our proposed task-specific knowledge. High-confidence data are identified when pseudo-labels from both sources align, while mismatches are treated as low-confidence data. A bimodal classifier is then trained to iteratively label the low-confidence data until a predefined criterion is met. We evaluate our SSL framework on emotion recognition and dementia detection tasks. Experimental results demonstrate that our method achieves competitive performance compared to fully supervised learning using only 30% of the labeled data and significantly outperforms two selected baselines.

半监督语音分析认知状态伪标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。