arXiv:2606.15984cs.CL2026-06

解决罗马尼亚语议会语音识别中的方言与人口偏见问题

ROMPAR: Morphological Completion and Demographic Unlearning for Romanian-Accented Speech Recognition

论文配图:ROMPAR: Morphological Completion and Demographic Unlearning for Romanian-Accented Speech Recognition
图 1 · 摘自论文原文
  • 多任务对抗训练确保年龄、性别、方言不变性
  • 词尾截断的形态补全使重构准确率达96.6% F1
  • 适合语音识别中需消除人口偏差的场景

自动转录议会发言面临人口偏见、方言差异及分割时的语句截断等挑战。本文提出17.80小时的罗马尼亚与摩尔多瓦议会语音语料库ROMPAR,包含双标注真实文本和重建词片段的显式标签。为构建鲁棒自动语音识别系统,我们提出一种多任务对抗训练框架,实现年龄、性别、方言上的群体不变性。通过引入对抗系数的指数衰减机制,缓解生成架构中对抗目标的不稳定性。此外,采用基于大语言模型的解码策略,结合位置依赖加权,促进截断词末的形态补全。实验表明,该框架显著降低词错误率(WER),在形态重构上达到96.6% F1分数。

原文摘要 · Abstract (English)

Automated transcription of parliamentary proceedings faces significant hurdles due to demographic bias, dialectal variation, and technical artifacts such as utterance truncation during segmentation. This paper introduces the ROManian PARliamentary Speech Corpus (ROMPAR) dataset, a 17.80-hour corpus of Romanian and Moldavian parliamentary speech, featuring double-annotated ground truth and explicit labels for reconstructed word fragments. To build a robust ASR system, we propose a multi-task adversarial training framework that enforces demographic invariance across age, gender, and dialect. We address the inherent instability of adversarial objectives in generative architectures by introducing an exponential decay mechanism for the adversarial coefficients. Furthermore, we implement an LLM-guided decoding strategy with position-dependent weighting to facilitate morphological completion of truncated terminal words. Our results demonstrate that the proposed framework significantly reduces WER and achieves an F1-score of 96.6% in morphological reconstruction.

语音识别形态补全对抗训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。