arXiv:2608.21176cs.SDcs.AI2026-08

通过定位语音失真位置,提升自动语音质量评估准确性。

DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization

论文配图:DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization
图 1 · 摘自论文原文
  • 引入失真定位作为辅助信息,指导质量评估模型关注关键失真区域。
  • 在多个公开数据集上超越现有方法,跨数据集泛化能力强。
  • 构建首个带帧级标注的局部失真语音数据集,支持精细分析。

自动语音质量评估旨在预测与人类主观感知一致的平均意见分(MOS),对语音生成、增强及通信系统评价至关重要。对于语音信号,尤其是合成语音,失真常局部出现,而整体感知质量通常由少数显著失真区域主导。然而,多数现有方法仅以语句级MOS进行优化,提供的是粗粒度监督,无法明确指示重要失真发生的位置。为此,本文首次引入显式失真定位作为辅助知识。我们构建了首个部分失真的语音数据集,包含帧级失真标注,并训练定位模型生成失真提示。基于这些提示,提出DAMOS框架,将失真定位信息融入MOS预测流程。在多个公开基准上的实验表明,DAMOS持续优于现有方法,且具备强跨数据集泛化能力,验证了显式失真定位在语音质量评估中的有效性。

原文摘要 · Abstract (English)

Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential for evaluating speech generation, enhancement, and communication systems. For speech signals, especially synthetic speech, distortions often occur locally, and overall perceptual quality is usually dominated by a small number of perceptually salient distortion regions. However, most existing methods are primarily optimized with utterance-level MOS, which provides only coarse-grained supervision and offer no explicit indication of where perceptually important distortions occur. To address this limitation, we introduce explicit distortion localization as auxiliary knowledge for speech quality assessment. We construct the first partially distorted speech dataset with frame-level distortion annotations and train a localization model to generate distortion cues. Building on these cues, we propose DAMOS, a distortion-aware speech quality assessment framework that integrates localization information into the MOS prediction pipeline. Experiments on multiple public benchmarks demonstrate that DAMOS consistently outperforms existing methods and exhibits strong cross-dataset generalization, validating the effectiveness of explicit distortion localization for speech quality assessment.

语音质量失真定位MOS预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。