用链式推理同时评估语音质量多个指标,更准更可靠。
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
- 通过自回归建模构建动态指标链,捕捉各评价指标间依赖关系。
- 在多种语音场景下,对PESQ、STOI等指标的预测精度显著超越基线。
- 支持单个或多个指标查询,且通过置信度优化减少误差传播,适合语音生成与降噪评估。
语音信号分析在语音质量评估与特征刻画任务中面临挑战,目标是同时预测多个感知与客观指标,如PESQ(感知语音质量评价)、STOI(短时客观可懂度)和MOS(平均意见分)。这些指标量纲不同、假设各异且相互依赖,联合估计难度大。为此,本文提出ARECHO(基于链式假设优化的自回归评估),一种基于自回归依赖建模的通用语音评估系统。其核心创新包括:(1) 全面的语音信息标记化流程;(2) 显式建模指标间依赖的动态分类器链;(3) 增强推理可靠性的两阶段置信度导向解码算法。实验表明,ARECHO在增强语音分析、语音生成评估及噪声语音评估等多种场景下均显著优于基线框架。其动态依赖建模提升了可解释性,支持子集查询(单个或多个指标),并通过置信度导向解码减少误差传播,实现无需参考的高效评估。
原文摘要 · Abstract (English)
Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and MOS (Mean Opinion Score) each capture different aspects of speech quality. However, these metrics often have different scales, assumptions, and dependencies, making joint estimation non-trivial. To address these issues, we introduce ARECHO (Autoregressive Evaluation via Chain-based Hypothesis Optimization), a chain-based, versatile evaluation system for speech assessment grounded in autoregressive dependency modeling. ARECHO is distinguished by three key innovations: (1) a comprehensive speech information tokenization pipeline; (2) a dynamic classifier chain that explicitly captures inter-metric dependencies; and (3) a two-step confidence-oriented decoding algorithm that enhances inference reliability. Experiments demonstrate that ARECHO significantly outperforms the baseline framework across diverse evaluation scenarios, including enhanced speech analysis, speech generation evaluation, and, noisy speech evaluation. Furthermore, its dynamic dependency modeling improves interpretability by capturing inter-metric relationships. Across tasks, ARECHO offers reference-free evaluation using its dynamic classifier chain to support subset queries (single or multiple metrics) and reduces error propagation via confidence-oriented decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。