挑战真实复杂语音退化,揭示轻量模型更优与评估指标缺陷。
The CCF AATC 2025 Speech Restoration Challenge: A Retrospective
- 设计统一模型应对声学干扰、编码压缩和增强算法引入的复合退化。
- 轻量判别模型(<10M参数)性能领先,生成模型在高信噪比下有重建偏差。
- 现有无参考指标与人评感知自然度负相关,需改进评估方法。
现实语音通信常面临多种退化的复合影响:声学干扰、编码压缩及上游增强算法引入的次生伪影。为弥合学术研究与真实场景的差距,我们发起CCF AATC 2025语音恢复挑战赛,目标是通用盲语音恢复,要求单一模型同时处理三类退化:声学退化、编码失真和二次处理伪影。本文全面回顾该挑战,详述数据集构建、任务设计,并系统分析25个参赛系统。报告三项关键发现:(1) 效率与规模权衡:顶尖系统表明,轻量判别架构(<10M参数)可实现最先进性能,兼顾恢复质量与部署约束;(2) 生成模型的权衡:生成与混合模型虽在理论感知指标上表现优异,但在高信噪比编码任务中存在“重建偏差”,在复杂次生伪影场景中易产生幻觉;(3) 评估指标鸿沟:秩相关分析显示,主流无参考指标(如DNSMOS)与人类主观评分(MOS)间存在强负相关(ρ=-0.8),表明当前指标可能过度奖励人工谱平滑而牺牲感知自然度。本文旨在为鲁棒语音恢复研究提供参考,并呼吁开发对生成伪影敏感的下一代评估指标。
原文摘要 · Abstract (English)
Real-world speech communication is rarely affected by a single type of degradation. Instead, it suffers from a complex interplay of acoustic interference, codec compression, and, increasingly, secondary artifacts introduced by upstream enhancement algorithms. To bridge the gap between academic research and these realistic scenarios, we introduced the CCF AATC 2025 Challenge. This challenge targets universal blind speech restoration, requiring a single model to handle three distinct distortion categories: acoustic degradation, codec distortion, and secondary processing artifacts. In this paper, we provide a comprehensive retrospective of the challenge, detailing the dataset construction, task design, and a systematic analysis of the 25 participating systems. We report three key findings that define the current state of the field: (1) Efficiency vs. Scale: Contrary to the trend of massive generative models, top-performing systems demonstrated that lightweight discriminative architectures (<10M parameters) can achieve state-of-the-art performance, balancing restoration quality with deployment constraints. (2) Generative Trade-off: While generative and hybrid models excel in theoretical perceptual metrics, breakdown analysis reveals they suffer from "reconstruction bias" in high-SNR codec tasks and struggle with hallucination in complex secondary artifact scenarios. (3) Metric Gap: Most critically, our rank correlation analysis exposes a strong negative correlation (\r{ho}=-0.8) between widely-used reference-free metrics (e.g., DNSMOS) and human MOS when evaluating hybrid systems. This indicates that current metrics may over-reward artificial spectral smoothness at the expense of perceptual naturalness. This paper aims to serve as a reference for future research in robust speech restoration and calls for the development of next-generation evaluation metrics sensitive to generative artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。