用模拟对话提升匈牙利语语音识别,更真实地还原说话人切换和停顿。
Speaker-Aware Simulation Improves Conversational Speech Recognition
- 将单说话人录音转为多说话人对话,加入时长相关的停顿建模。
- 在字符错误率上提升显著,尤其在匹配真实对话统计时效果更好。
- 适合低资源语言语音识别研究者,尤其关注对话建模的场景。
对话式语音识别(ASR)因缺乏大规模、高质量的多说话人对话数据及自然交互中的复杂时间动态而仍具挑战性。本文提出并实现了针对匈牙利语的说话人感知模拟对话(SASC)框架,并进一步提出改进版本C-SASC,通过基于话语时长的停顿建模,更真实地捕捉人类对话中的局部时间依赖关系,同时保持原方法的简洁与高效。我们利用BEA-Large语料生成合成匈牙利语对话,并与真实对话数据结合用于训练。在CallHome、BEA-Dialogue和GRASS等语料导出的对话统计下,对多种模拟配置进行广泛评估。实验表明,相比简单的拼接增强,说话人感知模拟对话始终带来性能提升;其中C-SASC在字符级错误率上实现系统性改善,但其有效性取决于源对话统计与目标领域的匹配程度。结果验证了该方法在匈牙利语ASR中的鲁棒性,也揭示了更精细时间建模在合成对话生成中的优势与局限。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) for conversational speech remains challenging due to the limited availability of large-scale, well-annotated multi-speaker dialogue data and the complex temporal dynamics of natural interactions. Speaker-aware simulated conversations (SASC) offer an effective data augmentation strategy by transforming single-speaker recordings into realistic multi-speaker dialogues. However, prior work has primarily focused on English data, leaving questions about the applicability to lower-resource languages. In this paper, we adapt and implement the SASC framework for Hungarian conversational ASR. We further propose C-SASC, an extended variant that incorporates pause modeling conditioned on utterance duration, enabling a more faithful representation of local temporal dependencies observed in human conversation while retaining the simplicity and efficiency of the original approach. We generate synthetic Hungarian dialogues from the BEA-Large corpus and combine them with real conversational data for ASR training. Both SASC and C-SASC are evaluated extensively under a wide range of simulation configurations, using conversational statistics derived from CallHome, BEA-Dialogue, and GRASS corpora. Experimental results show that speaker-aware conversational simulation consistently improves recognition performance over naive concatenation-based augmentation. While the additional duration conditioning in C-SASC yields modest but systematic gains--most notably in character-level error rates--its effectiveness depends on the match between source conversational statistics and the target domain. Overall, our findings confirm the robustness of speaker-aware conversational simulation for Hungarian ASR and highlight the benefits and limitations of increasingly detailed temporal modeling in synthetic dialogue generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。