arXiv:2510.05191cs.SDcs.AI2025-10被引 1

通过隐变量独立性实现可证明的语音属性转换

Provable Speech Attributes Conversion via Latent Independence

  • 设计非概率自编码器,强制隐变量与目标属性独立
  • 在说话人身份和情绪转换上均实现稳定可控
  • 首个具备理论保证的语音风格转换框架

尽管信号转换和解耦表示学习在音频、图像及多模态生成等领域展现潜力,但现有方法尤其是语音风格转换仍以经验为主,缺乏可靠的理论基础。本文提出一种通用语音属性转换框架,在合理假设下提供理论分析与保障。该框架基于非概率自编码器结构,引入隐变量与目标可控变量间的独立性约束,确保在给定风格变量条件下实现一致的信号变换,同时保留原始内容并修改指定属性。我们在说话人身份和情感等语音风格上验证了方法的通用性。定量评估证实了该方法的有效性和普适性。

原文摘要 · Abstract (English)

While signal conversion and disentangled representation learning have shown promise for manipulating data attributes across domains such as audio, image, and multimodal generation, existing approaches, especially for speech style conversion, are largely empirical and lack rigorous theoretical foundations to guarantee reliable and interpretable control. In this work, we propose a general framework for speech attribute conversion, accompanied by theoretical analysis and guarantees under reasonable assumptions. Our framework builds on a non-probabilistic autoencoder architecture with an independence constraint between the predicted latent variable and the target controllable variable. This design ensures a consistent signal transformation, conditioned on an observed style variable, while preserving the original content and modifying the desired attribute. We further demonstrate the versatility of our method by evaluating it on speech styles, including speaker identity and emotion. Quantitative evaluations confirm the effectiveness and generality of the proposed approach.

语音转换隐变量可证明性风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。