跨语言语音属性预测新方法,提升低资源语言表现
Cross-Language Speaker Attribute Prediction Using MIL and RL
- 用强化学习选关键语音片段,结合对抗域适应学通用特征
- 在5语言推特数据上性别预测宏F1提升显著,年龄预测也有改善
- 特别适合低资源语言的语音属性识别任务
我们研究在语言差异、领域不匹配和数据不平衡下的多语言说话人属性预测。提出RLMIL-DAT,一种基于强化学习的多实例学习框架扩展,结合基于强化学习的实例选择与域对抗训练,以促进语言无关的语音表示。在五个语言的推特语料库上进行少样本设置评估,在覆盖四十个语言的VoxCeleb2衍生语料库上进行零样本设置评估,目标为性别和年龄预测。在多种模型配置和多个随机种子下,RLMIL-DAT始终优于标准多实例学习和原始强化多实例学习框架。性别预测提升最明显,年龄预测仍具挑战但有小幅正向改进。消融实验表明,域对抗训练是性能提升的主要来源,通过抑制共享编码器中的语言特异性线索,实现从高资源英语到低资源语言的有效迁移。在较小的VoxCeleb2子集的零样本设置中,改进普遍为正但不够稳定,反映统计能力有限及向大量未见语言泛化的难度。总体表明,结合实例选择与对抗域适应是跨语言说话人属性预测有效且稳健的策略。
原文摘要 · Abstract (English)
We study multilingual speaker attribute prediction under linguistic variation, domain mismatch, and data imbalance across languages. We propose RLMIL-DAT, a multilingual extension of the reinforced multiple instance learning framework that combines reinforcement learning based instance selection with domain adversarial training to encourage language invariant utterance representations. We evaluate the approach on a five language Twitter corpus in a few shot setting and on a VoxCeleb2 derived corpus covering forty languages in a zero shot setting for gender and age prediction. Across a wide range of model configurations and multiple random seeds, RLMIL-DAT consistently improves Macro F1 compared to standard multiple instance learning and the original reinforced multiple instance learning framework. The largest gains are observed for gender prediction, while age prediction remains more challenging and shows smaller but positive improvements. Ablation experiments indicate that domain adversarial training is the primary contributor to the performance gains, enabling effective transfer from high resource English to lower resource languages by discouraging language specific cues in the shared encoder. In the zero shot setting on the smaller VoxCeleb2 subset, improvements are generally positive but less consistent, reflecting limited statistical power and the difficulty of generalizing to many unseen languages. Overall, the results demonstrate that combining instance selection with adversarial domain adaptation is an effective and robust strategy for cross lingual speaker attribute prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。