一次识别语音中多种说话风格,提升人机交互实用性
SpeechMLC: Speech Multi-label Classification
- 用Transformer解码器加交叉注意力提取多风格特征
- 通过语音生成模型增强数据,缓解标签不平衡问题
- 兼顾人类标注一致性影响,适合实际应用场景
本文提出一种多标签分类框架,用于在单个语音样本中检测多种说话风格。与以往聚焦单一风格的研究不同,该框架在统一结构中有效捕捉多种说话者特征,适用于通用人机交互场景。模型在Transformer解码器中引入交叉注意力机制,从输入语音中提取每个目标标签的显著特征。针对多标签语音数据集固有的数据不平衡问题,采用基于语音生成模型的数据增强技术。通过在已见和未见语料上的多项客观评估验证了模型有效性。此外,还分析了人类感知对分类精度的影响,考察了人工标注一致性的模型性能影响。
原文摘要 · Abstract (English)
In this paper, we propose a multi-label classification framework to detect multiple speaking styles in a speech sample. Unlike previous studies that have primarily focused on identifying a single target style, our framework effectively captures various speaker characteristics within a unified structure, making it suitable for generalized human-computer interaction applications. The proposed framework integrates cross-attention mechanisms within a transformer decoder to extract salient features associated with each target label from the input speech. To mitigate the data imbalance inherent in multi-label speech datasets, we employ a data augmentation technique based on a speech generation model. We validate our model's effectiveness through multiple objective evaluations on seen and unseen corpora. In addition, we provide an analysis of the influence of human perception on classification accuracy by considering the impact of human labeling agreement on model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。