arXiv:2512.04827cs.SDcs.LG2025-12

用可解释的体验契约评估语音与演唱质量,比传统评分更稳定可靠。

Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs

  • 以人类可理解的体验契约替代单一评分,构建服务图的质量审计框架。
  • 在多个数据集上验证,契约满意度比原始MOS更抗视角变换干扰。
  • 适合关注服务质量可解释性与跨系统对比的研究者和工程师。

主观平均意见分(MOS)仍是非侵入式语音与演唱质量评估的主流指标。然而,MOS是标量,会掩盖用户期望的异质性,忽略服务级目标,且难以在不同部署图间比较。本文提出一种契约驱动的用户体验审计框架:每个服务图G在一组人类可理解的体验契约C下评估,生成契约级满意度向量Q(G, C)。我们证明:(i) 经典MOS回归是退化契约集下的特例;(ii) 契约驱动的质量度量在图视图变换(如按系统或系统类型聚合)下比MOS更稳定;(iii) 学习契约的有效样本复杂度由契约语义决定,而非仅由契约维度决定。我们在URGENT2024 MOS(6.9k语音语段,含原始评分向量)和SingMOS v1(7,981个演唱片段;80个系统)上实例化该框架。在URGENT上,基于自监督WavLM嵌入训练契约感知神经审计器;在SingMOS上,使用公开评分向量与元数据进行契约驱动的图审计,无需解码音频。实证表明,审计器在MOS精度上媲美强基准模型,同时提供校准的契约概率;在SingMOS上,Q(G, C)的跨视图漂移显著小于原始MOS与纯图基线;在URGENT上,困难曲线显示,误设的“简单”契约反而比语义更匹配但结构更丰富的契约更难学习。

原文摘要 · Abstract (English)

Subjective mean opinion scores (MOS) remain the de-facto target for non-intrusive speech and singing quality assessment. However, MOS is a scalar that collapses heterogeneous user expectations, ignores service-level objectives, and is difficult to compare across deployment graphs. We propose a contract-driven QoE auditing framework: each service graph G is evaluated under a set of human-interpretable experience contracts C, yielding a contract-level satisfaction vector Q(G, C). We show that (i) classical MOS regression is a special case with a degenerate contract set, (ii) contract-driven quality is more stable than MOS under graph view transformations (e.g., pooling by system vs. by system type), and (iii) the effective sample complexity of learning contracts is governed by contract semantics rather than merely the dimensionality of C. We instantiate the framework on URGENT2024 MOS (6.9k speech utterances with raw rating vectors) and SingMOS v1 (7,981 singing clips; 80 systems). On URGENT, we train a contract-aware neural auditor on self-supervised WavLM embeddings; on SingMOS, we perform contract-driven graph auditing using released rating vectors and metadata without decoding audio. Empirically, our auditor matches strong MOS predictors in MOS accuracy while providing calibrated contract probabilities; on SingMOS, Q(G, C) exhibits substantially smaller cross-view drift than raw MOS and graph-only baselines; on URGENT, difficulty curves reveal that mis-specified "simple" contracts can be harder to learn than richer but better aligned contract sets.

质量评估契约审计语音处理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。