arXiv:2506.06071eess.AScs.CL2025-06中稿 · IEEE ASRU 2025被引 3

通过语音转换增强数据,让语音情感识别更公平

CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition

  • 基于置信度识别偏差样本并用语音转换生成新数据
  • 在多个模型上验证,显著降低不同人群的识别偏差
  • 无需修改模型或依赖人口属性,适合实际部署

语音情感识别系统中的偏见常源于说话人特征与情感标签之间的虚假关联,导致不同人口群体预测不公平。现有去偏方法多需修改模型结构或依赖人口标注,实用性受限。本文提出 CO-VADA——一种面向置信度的语音增强去偏方法,不改变模型架构且无需人口信息。该方法识别训练数据中反映偏差模式的样本,通过语音转换改变无关属性以生成新样本。这些增强样本引入与数据主导模式不同的说话人差异,引导模型关注情感相关特征。本框架兼容多种 SER 模型与语音转换工具,是提升 SER 公平性的可扩展、实用方案。

原文摘要 · Abstract (English)

Bias in speech emotion recognition (SER) systems often stems from spurious correlations between speaker characteristics and emotional labels, leading to unfair predictions across demographic groups. Many existing debiasing methods require model-specific changes or demographic annotations, limiting their practical use. We present CO-VADA, a Confidence-Oriented Voice Augmentation Debiasing Approach that mitigates bias without modifying model architecture or relying on demographic information. CO-VADA identifies training samples that reflect bias patterns present in the training data and then applies voice conversion to alter irrelevant attributes and generate samples. These augmented samples introduce speaker variations that differ from dominant patterns in the data, guiding the model to focus more on emotion-relevant features. Our framework is compatible with various SER models and voice conversion tools, making it a scalable and practical solution for improving fairness in SER systems.

语音识别情感分析公平性去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。