arXiv:2509.09212eess.AScs.SD2025-09

新评估方法精准分离音源分离中的失真与干扰,更贴近人耳感知。

MAPSS: Manifold-based Assessment of Perceptual Source Separation

  • 通过生成基础失真并用流形学习建模,分离失真与泄漏因素。
  • 在英、西语及音乐混合数据上,相关性超越18种主流指标。
  • 支持每秒75帧的精细分析,适合模型优化与主观评价验证。

音频源分离系统的客观评估仍难以匹配人类主观感知,尤其当竞争说话人干扰与目标信号失真同时存在时。本文提出感知分离(PS)与感知匹配(PM),一对互补度量,从设计上将泄漏与失真因素分离。采用侵入式方法,对混合中的参考波形生成剪裁、陷波滤波、音高偏移等基本失真。所有参考、失真及系统输出由预训练自监督模型独立编码,再通过扩散映射(diffusion maps)进行流形学习,使欧氏距离与编码表示的不相似性对齐。在此流形上,PM通过输出到其参考及对应失真的距离衡量源的自失真;PS则额外考虑输出到非归属参考与失真的距离,以捕捉泄漏。两项度量均具有可微性,分辨率高达每秒75帧,支持精细化优化与分析。进一步推导出两种度量的帧级确定性误差半径及非渐近高概率置信区间。在英语、西班牙语和音乐混合数据上的实验表明,相较于18种广泛使用的度量,PS与PM在与主观平均意见分(MOS)的线性与排名相关性上几乎始终位列第一或第二。

原文摘要 · Abstract (English)

Objective assessment of audio source-separation systems still mismatches subjective human perception, especially when interference from competing talkers and distortion of the target signal interact. We introduce Perceptual Separation (PS) and Perceptual Match (PM), a complementary pair of measures that, by design, isolate these leakage and distortion factors. Our intrusive approach generates a set of fundamental distortions, e.g., clipping, notch filter, and pitch shift from each reference waveform signal in the mixture. Distortions, references, and system outputs from all sources are independently encoded by a pre-trained self-supervised model, then aggregated and embedded with a manifold learning technique called diffusion maps, which aligns Euclidean distances on the manifold with dissimilarities of the encoded waveform representations. On this manifold, PM captures the self-distortion of a source by measuring distances from its output to its reference and associated distortions, while PS captures leakage by also accounting for distances from the output to non-attributed references and distortions. Both measures are differentiable and operate at a resolution as high as 75 frames per second, allowing granular optimization and analysis. We further derive, for both measures, frame-level deterministic error radius and non-asymptotic, high-probability confidence intervals. Experiments on English, Spanish, and music mixtures show that, against 18 widely used measures, the PS and PM are almost always placed first or second in linear and rank correlations with subjective human mean-opinion scores.

音频分离主观评估流形学习可微度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。