构建首个多模态多语言语音生成提示库,提升真实场景下的语音合成效果。
$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation
- 采用多智能体标注框架,实现跨模态的精细标签体系。
- 在真实场景下显著提升语音生成质量,尤其在低质提示条件下。
- 适合关注实际应用、跨语言语音合成的研究者使用。
零样本语音生成技术已能根据语音提示合成模仿说话人身份和语调的语音,但在真实场景中,高质量语音提示往往缺失、不完整或超出领域范围,导致模型表现受限。这主要源于训练数据与推理时提示语音之间的质量差异。为此,我们提出M³PDB,首个大规模、多模态、多标签、多语言的语音生成提示数据库,旨在支持鲁棒的提示选择。数据构建采用创新的多模态多智能体标注框架,实现跨模态的精确、分层标注。同时,我们设计了一种轻量高效的提示选择策略,适用于实时、资源受限的推理环境。实验表明,该数据库与策略可有效支持多种挑战性语音生成场景。我们希望推动社区从标准基准性能优化转向更真实、多样化的应用场景研究。代码与数据集开源:https://github.com/hizening/M3PDB。
原文摘要 · Abstract (English)
Recent advancements in zero-shot speech generation have enabled models to synthesize speech that mimics speaker identity and speaking style from speech prompts. However, these models' effectiveness is significantly limited in real-world scenarios where high-quality speech prompts are absent, incomplete, or out of domain. This issue arises primarily from a significant quality mismatch between the speech data utilized for model training and the input prompt speech during inference. To address this, we introduce $\text{M}^3\text{PDB}$, the first large-scale, multi-modal, multi-label, and multilingual prompt database designed for robust prompt selection in speech generation. Our dataset construction leverages a novel multi-modal, multi-agent annotation framework, enabling precise and hierarchical labeling across diverse modalities. Furthermore, we propose a lightweight yet effective prompt selection strategy tailored for real-time, resource-constrained inference settings. Experimental results demonstrate that our proposed database and selection strategy effectively support various challenging speech generation scenarios. We hope our work can inspire the community to shift focus from improving performance on standard benchmarks to addressing more realistic and diverse application scenarios in speech generation. Code and dataset are available at: https://github.com/hizening/M3PDB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。