用大模型实现少样本文本归一化,降低人工规则依赖。
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech
- 基于提示工程的LLM方法,无需手工规则
- 八种语言下词错误率显著低于现有系统
- 适合多语言、低资源场景的文本转语音研究
文本归一化(TN)是文本转语音(TTS)系统的关键预处理步骤,将书面形式转换为标准口语表达。传统TN系统虽准确率高,但需大量工程投入,难以扩展,且在低资源语言中覆盖困难。本文提出PolyNorm,一种基于大语言模型(LLM)的提示式归一化方法,旨在减少对人工规则的依赖,并实现更广的语言适用性,仅需极少人工干预。此外,我们设计了一套语言无关的自动数据清洗与评估流程,支持跨语言大规模实验。在八种语言上的实验表明,相比生产级系统,词错误率(WER)持续降低。为促进后续研究,我们发布了PolyNorm-Benchmark,一个涵盖多种文本归一化现象的多语言数据集。
原文摘要 · Abstract (English)
Text Normalization (TN) is a key preprocessing step in Text-to-Speech (TTS) systems, converting written forms into their canonical spoken equivalents. Traditional TN systems can exhibit high accuracy, but involve substantial engineering effort, are difficult to scale, and pose challenges to language coverage, particularly in low-resource settings. We propose PolyNorm, a prompt-based approach to TN using Large Language Models (LLMs), aiming to reduce the reliance on manually crafted rules and enable broader linguistic applicability with minimal human intervention. Additionally, we present a language-agnostic pipeline for automatic data curation and evaluation, designed to facilitate scalable experimentation across diverse languages. Experiments across eight languages show consistent reductions in the word error rate (WER) compared to a production-grade-based system. To support further research, we release PolyNorm-Benchmark, a multilingual data set covering a diverse range of text normalization phenomena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。