让语音更讨喜:通过自动评分实现语音风格可控转换
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
- 用预训练评分模型自动标注海量语音数据的讨喜度
- 转换后语音讨喜度提升,同时保留说话人身份和语义内容
- 适合语音合成、人机交互等需要情感适配的场景
语音的感知讨喜度在伴侣选择、广告传播等社交互动中至关重要。若能提供针对目标受众定制的讨喜语音样本,用户便可调整自身发声风格与音质,促进沟通顺畅。为此,本文提出一种语音转换方法,可在保持说话人身份和语言内容不变的前提下,控制输入语音的讨喜度。为提升训练数据可扩展性,我们基于现有语音讨喜度数据集训练了一个讨喜度预测器,并用于自动标注大规模语音合成语料库的讨喜度标签。实验表明,该预测器输出与人工评分具有显著相关性。主观与客观评估进一步验证了所提方法能有效调控语音讨喜度,同时保持说话人身份和语言内容的一致性。
原文摘要 · Abstract (English)
Perceived voice likability plays a crucial role in various social interactions, such as partner selection and advertising. A system that provides reference likable voice samples tailored to target audiences would enable users to adjust their speaking style and voice quality, facilitating smoother communication. To this end, we propose a voice conversion method that controls the likability of input speech while preserving both speaker identity and linguistic content. To improve training data scalability, we train a likability predictor on an existing voice likability dataset and employ it to automatically annotate a large speech synthesis corpus with likability ratings. Experimental evaluations reveal a significant correlation between the predictor's outputs and human-provided likability ratings. Subjective and objective evaluations further demonstrate that the proposed approach effectively controls voice likability while preserving both speaker identity and linguistic content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。