用嵌入空间分类器发现合成语音超分辨率存在分布偏差。
Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers
- 用多种音频嵌入训练线性分类器区分真实与合成音频。
- 即使感知质量高、指标优秀,合成音频仍可被几乎100%识别出。
- 揭示了当前模型在分布拟合上的根本缺陷,适合模型评估者阅读。
生成对抗网络(GANs)和扩散模型近期在音频超分辨率(ADSR)任务中达到顶尖性能,能从窄带输入生成听觉上逼真的宽带音频。然而现有评估主要依赖信号级或感知指标,未解决合成超分辨率音频与真实宽带音频分布匹配程度的问题。本文通过分析不同嵌入空间中真实与合成音频的可分性来解决该问题。针对语音和音乐,在中频段(4→16 kHz)和全频段(16→48 kHz)上进行上采样任务,利用多种音频嵌入训练线性分类器以区分真实与合成样本。对比客观指标和主观听感测试发现,嵌入分类器几乎能完美分离两类音频,即使生成音频在感知质量上表现优异且取得领先指标分数。该现象在多个数据集和模型(包括最新扩散模型)中一致出现,凸显了当前ADSR模型在感知质量与真实分布保真度之间的持续差距。
原文摘要 · Abstract (English)
Generative adversarial networks (GANs) and diffusion models have recently achieved state-of-the-art performance in audio super-resolution (ADSR), producing perceptually convincing wideband audio from narrowband inputs. However, existing evaluations primarily rely on signal-level or perceptual metrics, leaving open the question of how closely the distributions of synthetic super-resolved and real wideband audio match. Here we address this problem by analyzing the separability of real and super-resolved audio in various embedding spaces. We consider both middle-band ($4\to 16$~kHz) and full-band ($16\to 48$~kHz) upsampling tasks for speech and music, training linear classifiers to distinguish real from synthetic samples based on multiple types of audio embeddings. Comparisons with objective metrics and subjective listening tests reveal that embedding-based classifiers achieve near-perfect separation, even when the generated audio attains high perceptual quality and state-of-the-art metric scores. This behavior is consistent across datasets and models, including recent diffusion-based approaches, highlighting a persistent gap between perceptual quality and true distributional fidelity in ADSR models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。