评估语音生成中属性控制的稳定性,发现常见模型会意外改变非目标特征。
Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation

- 通过对比实验分析三类语音生成模型的属性控制能力
- 85%以上输出在目标特征变化时伴随显著非目标特征偏移
- 提出无需训练的候选重排序方法,提升多属性稳定生成率
自然语言描述已成为控制生成语音的灵活接口。现有评估主要关注输出是否匹配提示,但提示匹配本身无法揭示非目标特征是否保持稳定。我们通过一项包含5,940个输出的受控配对审计,考察了CosyVoice3、VoxCPM2和Fish-Speech-S2三个语音生成系统的表现。实验覆盖六位参考说话人、十段文本、三个随机种子和十一组条件。基于声学、韵律、内容和说话人特征测量,发现多数目标方向响应伴随着超出描述特定信号集的变化。这一现象在目标响应超过基线种子变异的情况下依然存在,且不同系统间差异显著。我们进一步提出VoDER-Cal,一种无需训练的候选选择器,在保留强目标响应的同时偏好更小的非目标偏差。三候选池将联合成功率从单样本直接生成的4.8%提升至约14%。在匹配三候选预算内,VoDER-Cal将保留测试集的非目标偏差从0.344降至0.276,并提升听者评分的保真度。因此,保真度敏感评估可补充提示遵循性评估,候选重排序则提供了实用的推理时改进方案。代码、配置文件与分析脚本已公开于https://github.com/intelland/VoDER。
原文摘要 · Abstract (English)
Natural-language descriptions have become a flexible interface for controlling generated speech. Existing evaluations largely assess whether an output matches a prompt, but prompt matching alone does not reveal whether characteristics outside the intended change remain stable. We examine this distinction through a controlled paired audit of three speech-generation systems: CosyVoice3, VoxCPM2, and Fish-Speech-S2. The evaluation contains 5,940 outputs spanning six reference speakers, ten texts, three random seeds, and eleven conditions. Using acoustic, prosodic, content, and speaker measurements, we find that responses in the expected target direction are frequently accompanied by changes outside descriptor-specific signal-level target sets. This pattern remains among outputs whose target response exceeds baseline seed variation, and the accompanying changes differ substantially across systems. We further introduce VoDER-Cal, a training-free candidate selector that retains sufficiently strong target responses while favoring smaller off-target deviations. A three-candidate pool raises the joint success rate from 4.8% under single-sample direct generation to approximately 14% for all candidate-selection policies. Within the matched three-candidate budget, VoDER-Cal reduces held-out off-target deviation from 0.344 under target-only selection to 0.276 and improves listener-rated preservation. Preservation-sensitive evaluation therefore complements prompt-adherence evaluation, while candidate reranking offers a practical inference-time improvement. Code, configuration files, and analysis scripts are available at https://github.com/intelland/VoDER
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。