arXiv:2505.17093eess.AScs.CL2025-05被引 1

将人物描述自动转为语音属性,实现可控语音合成。

P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech

  • 通过两种策略从人物描述生成语音属性,支持结构化与风格化输入。
  • 相比基线降低5%词错误率,语音质量评分提升0.33分。
  • 首次建立人物描述与语音合成的关联,揭示大模型中的社会偏见。

尽管基于角色的大语言模型和基于提示的文本转语音(TTS)系统已取得显著进展,但用户在尝试根据隐式人物描述生成符合期望的语音时仍面临可用性挑战。大多数用户缺乏专业知识来精确指定语音属性,导致TTS系统常误解其需求。为此,我们提出首个可从人物描述自动生成语音属性的框架P2VA。该方法采用两种策略:面向结构化语音属性的P2VA-C,以及面向丰富风格描述的P2VA-O。评估表明,P2VA-C使词错误率(WER)降低5%,语音质量评分(MOS)提升0.33分。据我们所知,P2VA是首个建立人物描述与语音合成之间联系的框架。此外,我们发现当前大语言模型在转换过程中会引入社会偏见。实验与分析进一步揭示了构建角色-语音系统的挑战。

原文摘要 · Abstract (English)

While persona-driven large language models (LLMs) and prompt-based text-to-speech (TTS) systems have advanced significantly, a usability gap arises when users attempt to generate voices matching their desired personas from implicit descriptions. Most users lack specialized knowledge to specify detailed voice attributes, which often leads TTS systems to misinterpret their expectations. To address these gaps, we introduce Persona-to-Voice-Attribute (P2VA), the first framework enabling voice generation automatically from persona descriptions. Our approach employs two strategies: P2VA-C for structured voice attributes, and P2VA-O for richer style descriptions. Evaluation shows our P2VA-C reduces WER by 5% and improves MOS by 0.33 points. To the best of our knowledge, P2VA is the first framework to establish a connection between persona and voice synthesis. In addition, we discover that current LLMs embed societal biases in voice attributes during the conversion process. Our experiments and findings further provide insights into the challenges of building persona-voice systems.

语音合成角色驱动大模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。