arXiv:2410.05698cs.CLcs.AI2024-10EMNLP

用两步法解决法语发音学习数据少的难题

A Two-Step Approach for Data-Efficient French Pronunciation Learning

  • 分两步完成音素转换与词后处理,降低对标注数据依赖
  • 仅用少量句子级发音数据就实现有效建模
  • 适合资源有限环境下法语语音学习研究者使用

近期研究在处理法语复杂音韵现象时,依赖大量语言学知识或句级发音数据,但这些资源的构建成本高且困难。为此,本文提出一种新颖的两步法,包括音位转换和词后处理两个任务,并在仅有少量句级发音数据的条件下验证其有效性。结果表明,该方法能有效缓解大规模标注数据不足的问题,在资源受限环境下仍可实现对法语音韵现象的可靠建模,为低资源法语发音学习提供可行方案。

原文摘要 · Abstract (English)

Recent studies have addressed intricate phonological phenomena in French, relying on either extensive linguistic knowledge or a significant amount of sentence-level pronunciation data. However, creating such resources is expensive and non-trivial. To this end, we propose a novel two-step approach that encompasses two pronunciation tasks: grapheme-to-phoneme and post-lexical processing. We then investigate the efficacy of the proposed approach with a notably limited amount of sentence-level pronunciation data. Our findings demonstrate that the proposed two-step approach effectively mitigates the lack of extensive labeled data, and serves as a feasible solution for addressing French phonological phenomena even under resource-constrained environments.

语音合成法语低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。