用大模型生成语音断句标注数据,降低人工成本。
Synthetic Data Generation for Phrase Break Prediction with Large Language Model
- 用大模型自动生成断句标注,替代人工标注。
- 跨语言验证显示生成数据有效提升预测性能。
- 适合语音合成、数据稀缺任务的研究者参考。
当前语音合成中的断句预测方法虽关键,但严重依赖大量音频或文本的人工标注,成本高昂。语音领域的固有变异性(受语音因素影响)也使高质量数据获取困难。近期大语言模型(LLM)在NLP中成功用于生成定制化合成数据,减少人工标注需求。受此启发,我们探索利用LLM生成合成断句标注,通过与传统标注对比,并在多语言上评估其有效性,以应对人工标注和语音任务的数据挑战。结果表明,基于LLM的合成数据生成能有效缓解断句预测中的数据难题,凸显了LLM在语音领域作为可行解决方案的潜力。
原文摘要 · Abstract (English)
Current approaches to phrase break prediction address crucial prosodic aspects of text-to-speech systems but heavily rely on vast human annotations from audio or text, incurring significant manual effort and cost. Inherent variability in the speech domain, driven by phonetic factors, further complicates acquiring consistent, high-quality data. Recently, large language models (LLMs) have shown success in addressing data challenges in NLP by generating tailored synthetic data while reducing manual annotation needs. Motivated by this, we explore leveraging LLM to generate synthetic phrase break annotations, addressing the challenges of both manual annotation and speech-related tasks by comparing with traditional annotations and assessing effectiveness across multiple languages. Our findings suggest that LLM-based synthetic data generation effectively mitigates data challenges in phrase break prediction and highlights the potential of LLMs as a viable solution for the speech domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。