利用标注者相似性,实现新标注者零成本情感识别个性化。
More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition
- 通过预训练模型匹配相似标注者,减少新标注者数据依赖。
- 仅需少量标注数据即可实现对未见标注者的精准预测。
- 适合需要快速适配新标注者的真实场景部署。
语音情感识别系统通常基于多个标注者的评分生成共识值,但难以预测单个标注者的标注结果。现有模型虽可学习所有标注者行为,但面对新标注者时需大量标注数据进行适应。本文提出利用标注者间的相似性:通过在大规模标注者群体上预训练的模型,识别出与新标注者相似的已有标注者。仅需少量新标注者标注数据,即可基于相似标注者进行预测,实现无需训练的标注结果迁移,支持低资源个性化。实验表明,该方法显著优于其他无需训练的基线方法,为轻量级情感适应提供可行路径,适用于真实世界部署。
原文摘要 · Abstract (English)
Speech emotion recognition systems often predict a consensus value generated from the ratings of multiple annotators. However, these models have limited ability to predict the annotation of any one person. Alternatively, models can learn to predict the annotations of all annotators. Adapting such models to new annotators is difficult as new annotators must individually provide sufficient labeled training data. We propose to leverage inter-annotator similarity by using a model pre-trained on a large annotator population to identify a similar, previously seen annotator. Given a new, previously unseen, annotator and limited enrollment data, we can make predictions for a similar annotator, enabling off-the-shelf annotation of unseen data in target datasets, providing a mechanism for extremely low-cost personalization. We demonstrate our approach significantly outperforms other off-the-shelf approaches, paving the way for lightweight emotion adaptation, practical for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。