arXiv:2602.03245eess.AScs.CL2026-02中稿 · and presented at J…被引 1

将《小王子》译成克罗地亚查卡维安方言,打造可直接用于AI训练的音文对齐数据集。

Mići Princ -- A Little Boy Teaching Speech Technologies the Chakavian Dialect

  • 将小说文本与语音逐字对齐,构建结构化多模态数据集。
  • 用Whisper模型适配方言后,词错误率降低一半,字符错误率减少三分之二。
  • 适合方言语音识别、文化遗产数字化及跨语言AI研究者使用。

本文发布《小王子》克罗地亚查卡维安方言版的印刷本与音频书,构建了文本与语音在每个字词层面完全对齐的计算机可读数据集。该数据集已上传至CLARIN.SI资源库,旨在永久保存这一珍贵方言内容,并支持人工智能应用。我们以Whisper-large-v3模型为例,将其从标准克罗地亚语适配至查卡维安方言,测试结果显示词错误率下降50%,字符错误率最高减少67%。该数据集不仅可用于语音识别研究,还可推动方言保护与数字人文发展,未来有望转化为在线数字版本,让更广泛人群感受这部经典在方言中的独特魅力。

原文摘要 · Abstract (English)

This paper documents our efforts in releasing the printed and audio book of the translation of the famous novel The Little Prince into the Chakavian dialect, as a computer-readable, AI-ready dataset, with the textual and the audio components of the two releases now aligned on the level of each written and spoken word. Our motivation for working on this release is multiple. The first one is our wish to preserve the highly valuable and specific content beyond the small editions of the printed and the audio book. With the dataset published in the CLARIN.SI repository, this content is from now on at the fingertips of any interested individual. The second motivation is to make the data available for various artificial-intelligence-related usage scenarios, such as the one we follow upon inside this paper already -- adapting the Whisper-large-v3 open automatic speech recognition model, with decent performance on standard Croatian, to Chakavian dialectal speech. We can happily report that with adapting the model, the word error rate on the selected test data has being reduced to a half, while we managed to remove up to two thirds of the error on character level. We envision many more usages of this dataset beyond the set of experiments we have already performed, both on tasks of artificial intelligence research and application, as well as dialectal research. The third motivation for this release is our hope that this, now highly structured dataset, will be transformed into a digital online edition of this work, allowing individuals beyond the research and technology communities to enjoy the beauty of the message of the little boy in the desert, told through the spectacular prism of the Chakavian dialect.

语音识别方言保护多模态数据AI训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。