提出新方法,让语音质量评估能实时逐步更新。
ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling

- 将语音质量评估改为分层自回归建模,逐步细化判断。
- 2秒前缀下误差降低48%,有效感知上下文约4-6秒。
- 适合流式语音系统、生成模型的质量监控场景。
语音质量通常在完整语句上评估,但流式和生成系统需从部分音频中逐步推断。现有模型依赖完整上下文,在输入受限时性能下降。本文在ARECHO基础上提出ANCHOR,将增量评估重新建模为多分辨率自回归任务。通过双分辨率标记与层次化结构,单个解码器同时建模片段级与语句级质量,实现从粗到细的逐步优化。实验表明,该方法在部分输入下表现显著稳健,2秒前缀下PLCMOS误差减少48%。收敛分析揭示有效感知上下文约为4-6秒。压力测试进一步识别出局部损坏下的结构化外推偏差。结果证明,层次监督能提升增量预测能力,并阐明感知质量随时间累积的机制。
原文摘要 · Abstract (English)
While speech quality is typically assessed on complete utterances, streaming and generative systems require incremental estimation from partial audio. Existing predictors assume full context, degrading on prefix-constrained inputs. Extending ARECHO, we propose ANCHOR, reformulating incremental assessment as a multi-resolution autoregressive task. It models chunk- and utterance-level quality within a single decoder using dual-resolution tokens and a resolution-aware hierarchy for coarse-to-fine refinement. Experiments show substantial robustness under partial input, including a 48% PLCMOS error reduction on 2-second prefixes. Convergence analysis reveals a 4-6 s effective perceptual context horizon. A stress test further isolates structured extrapolation biases under localized corruption. Results demonstrate that hierarchical supervision improves incremental prediction and elucidates how perceptual quality accumulates over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。