提出新指标U3D,更准确评估语音合成中的动态节奏相似性
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
- 用U3D分析语音动态节奏模式,突破传统嵌入只关注音色等静态特征的局限
- 发现现有ASV嵌入忽略节奏等动态特征,导致语音克隆相似性评估失准
- 适合语音合成、声纹识别领域研究者,尤其关注身份一致性评估的场景
由于语音身份具有多维度特性,建模极具挑战。当前生成式语音系统常使用为区分目的设计的自动声纹验证(ASV)嵌入来评估身份,但这类嵌入主要捕捉静态特征如音色和音高范围,忽视了节奏等动态特征。本文揭示了影响说话人相似性测量的混杂因素,并提出新度量方法U3D,专门用于评估说话人动态节奏模式的一致性。该工作推动了在日益强大的语音克隆系统背景下,对说话人身份一致性的评估进展。代码已公开。
原文摘要 · Abstract (English)
Modeling voice identity is challenging due to its multifaceted nature. In generative speech systems, identity is often assessed using automatic speaker verification (ASV) embeddings, designed for discrimination rather than characterizing identity. This paper investigates which aspects of a voice are captured in such representations. We find that widely used ASV embeddings focus mainly on static features like timbre and pitch range, while neglecting dynamic elements such as rhythm. We also identify confounding factors that compromise speaker similarity measurements and suggest mitigation strategies. To address these gaps, we propose U3D, a metric that evaluates speakers' dynamic rhythm patterns. This work contributes to the ongoing challenge of assessing speaker identity consistency in the context of ever-better voice cloning systems. We publicly release our code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。