用高资源语言数据提升低资源语言语音情绪识别效果
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
- 通过语音到语音翻译生成目标语言的带标注数据
- 在多个语言和模型上验证了方法的有效性
- 适合缺乏标注数据的多语言情绪识别研究者
语音情绪识别(SER)是实现通用人工智能代理自然人机交互的关键。然而,由于除英语和中文外其他语言缺乏标注数据,构建鲁棒的多语言SER系统仍具挑战。本文提出一种利用高资源语言数据提升低资源语言SER性能的方法。具体而言,我们采用富有表现力的语音到语音翻译(S2ST)结合新型自举式数据选择流程,在目标语言中生成带标注数据。大量实验表明,该方法在不同上游模型和语言间均具有效性和泛化能力。结果表明,该方法可促进更可扩展、更鲁棒的多语言SER系统发展。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) is a crucial component in developing general-purpose AI agents capable of natural human-computer interaction. However, building robust multilingual SER systems remains challenging due to the scarcity of labeled data in languages other than English and Chinese. In this paper, we propose an approach to enhance SER performance in low SER resource languages by leveraging data from high-resource languages. Specifically, we employ expressive Speech-to-Speech translation (S2ST) combined with a novel bootstrapping data selection pipeline to generate labeled data in the target language. Extensive experiments demonstrate that our method is both effective and generalizable across different upstream models and languages. Our results suggest that this approach can facilitate the development of more scalable and robust multilingual SER systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。