为中文语音合成设计可解释的细粒度诊断框架,解决模型质量评估难题。
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

- 构建12维感知评价体系,覆盖稳定性到表现力
- 用对抗扰动和专家样本生成高质量诊断数据集
- 模型输出带推理过程的评分,适合研究者调试语音合成
尽管生成式文本转语音(TTS)模型已接近人类水平,但传统整体指标无法诊断细微声学缺陷或解释感知性能下降。为此,我们提出针对中文的TTS-PRISM多维诊断框架:首先建立涵盖稳定性到高级表现力的12维评价体系;其次设计基于对抗扰动与专家锚点的目标合成管道,构建高质量诊断数据集;最后通过架构驱动的指令微调,将明确评分标准与推理逻辑嵌入高效端到端模型。在包含1,600个样本的黄金测试集上,TTS-PRISM展现出更强的人类对齐能力。对六种TTS范式的分析揭示了细粒度能力差异的直观诊断标志。代码与检查点已在GitHub开源。
原文摘要 · Abstract (English)
While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explain perceptual collapse. To address this, we propose TTS-PRISM, a multi-dimensional diagnostic framework for Mandarin. First, we establish a 12-dimensional schema spanning stability to advanced expressiveness. Second, we design a targeted synthesis pipeline with adversarial perturbations and expert anchors to build a high-quality diagnostic dataset. Third, schema-driven instruction tuning embeds explicit scoring criteria and reasoning into an efficient end-to-end model. Experiments on a 1,600-sample Gold Test Set show TTS-PRISM outperforms generalist models in human alignment. Profiling six TTS paradigms establishes intuitive diagnostic flags that reveal fine-grained capability differences. TTS-PRISM is open-source, with code and checkpoints at https://github.com/xiaomi-research/tts-prism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。