arXiv:2509.16781cs.CL2025-09EMNLP被引 1

构建了最大规模的罗马尼亚语方言语音数据集,用于识别地域差异。

MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language

  • 用多目标对抗训练让模型区分方言,同时忽略性别年龄等次要特征。
  • Wav2Vec2-Large在方言与年龄双重对抗下,方言识别准确率达93.08%。
  • 适合研究方言识别、语音不变性建模的研究者和开发者。

本文提出了MoRoVoc,目前规模最大、用于分析口语罗马尼亚语地域差异的数据集,包含超过93小时音频和88,192个音频样本,覆盖罗马尼亚与摩尔多瓦两地的口语。我们进一步提出一种多目标对抗训练框架,将说话人年龄和性别作为对抗目标,使模型在主任务上具备判别能力的同时,对这些次要属性保持不变性。对抗系数通过元学习动态调整以优化性能。实验表明,使用性别作为对抗目标时,Wav2Vec2-Base在方言识别任务上达到78.21%准确率;当同时以方言和年龄为对抗目标时,Wav2Vec2-Large在性别分类任务中准确率达93.08%。

原文摘要 · Abstract (English)

This paper introduces MoRoVoc, the largest dataset for analyzing the regional variation of spoken Romanian. It has more than 93 hours of audio and 88,192 audio samples, balanced between the Romanian language spoken in Romania and the Republic of Moldova. We further propose a multi-target adversarial training framework for speech models that incorporates demographic attributes (i.e., age and gender of the speakers) as adversarial targets, making models discriminative for primary tasks while remaining invariant to secondary attributes. The adversarial coefficients are dynamically adjusted via meta-learning to optimize performance. Our approach yields notable gains: Wav2Vec2-Base achieves 78.21% accuracy for the variation identification of spoken Romanian using gender as an adversarial target, while Wav2Vec2-Large reaches 93.08% accuracy for gender classification when employing both dialect and age as adversarial objectives.

方言识别语音数据集对抗训练罗马尼亚语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。