让语音和文字描述在共享空间对齐,支持更丰富的语音风格建模。
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
- 双编码器将语音与风格文本映射到同一向量空间。
- 在语音风格检索等任务上超越现有模型,尤其在复合风格评估中表现优异。
- 适合需要精细控制语音风格的语音合成与内容生成研究者。
我们提出ParaSpeechCLAP,一类双编码器模型,将语音与风格文本描述映射到共享嵌入空间,支持丰富且多维度的内在(说话人级)与情境(语句级)描述符,如音高、质感与情感,远超现有模型处理范围。我们分别训练了内在、情境及统一联合模型,发现专用模型在单一风格维度上更强,而联合模型在组合式评估中表现更优。此外,我们证明了在内在模型中加入分类损失和类别平衡训练可提升性能。实验表明,ParaSpeechCLAP在风格描述检索、语音属性分类以及作为推理阶段风格提示语音合成的奖励模型方面均优于基线。模型与代码已开源。
原文摘要 · Abstract (English)
We introduce ParaSpeechCLAP, a family of dual-encoder models that map speech and text style captions into a shared embedding space, supporting rich intrinsic (speaker-level) and situational (utterance-level) descriptors, such as pitch, texture, and emotion, beyond the narrow set handled by existing models. We train separate Intrinsic and Situational models alongside a unified Combined model, finding that specialized models are stronger on individual style dimensions while the unified model excels on compositional evaluation. We further show that ParaSpeechCLAP-Intrinsic benefits from an additional classification loss and class-balanced training. We demonstrate performance on style caption retrieval, speech attribute classification, and usability as inference-time reward models for style-prompted TTS. ParaSpeechCLAP models outperform baselines on most metrics across all three applications. Our models and code are released at https://github.com/ajd12342/paraspeechclap .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。