实测发现语音增强效果因任务而异,不能一概而论。
Navigating Speech Enhancement for Real-Time MRI: A Systematic Assessment of Signal Quality, Source Preservation, and Downstream Tasks
- 对比三种现成增强系统在五大数据集上的表现
- 高评分不等于更好识别率或保真度,结果依赖下游任务
- 无系统在所有场景下最优,需按任务选择
实时磁共振成像(rtMRI)录音受扫描仪噪声严重干扰,但通用语音增强是否提升语音研究与下游处理仍不明确。本文评估了三种现成系统——Denoiser、PASE 和 RE-USE——在五个 rtMRI 数据集上的表现,采用自然输入、干净输入探针和存档的加性噪声配对探针。多任务评估涵盖学习质量预测器、说话人与音素表征、基于参考的可懂度与质量指标、声学-音位探针、自动语音识别(ASR)及副语言任务。核心发现为:增强效果具有任务依赖性——预测质量得分更高并不意味着更优的 ASR 表现或更强的源信号保真度。在 15 组数据集-识别器组合中,RE-USE 在 11 组中获得更低的词错误率估计值,而 Denoiser 在 13 组中表现更差。在配对加性噪声探针中,PASE 与 RE-USE 提升了识别音素一致性、可懂度与感知质量估计值。Denoiser 提升了识别音素一致性和短时客观可懂度(STOI),但降低了说话人嵌入相似性。无任一系统在所有数据集、识别器与评价指标上均最优。因此,增强后的 rtMRI 音频应视为特定任务的转换产物,而非原音频或传统处理波形的普遍优化替代品。
原文摘要 · Abstract (English)
Audio recorded during real-time magnetic resonance imaging (rtMRI) is heavily contaminated by scanner noise, but it remains unclear whether general-purpose speech enhancement improves the signal for speech research and downstream processing. Three off-the-shelf systems---Denoiser, PASE, and RE-USE---are evaluated across five rtMRI corpora using naturally recorded inputs, a clean-input probe, and an archived paired additive-noise probe. The multi-task evaluation spans learned quality predictors, speaker and phone representations, reference-based intelligibility and quality measures, acoustic--phonetic probes, automatic speech recognition (ASR), and paralinguistic tasks. The central result is that enhancement effects are endpoint dependent: higher predicted-quality scores do not reliably imply better ASR performance or greater source fidelity. Across 15 corpus--recognizer comparisons using corpus-provided processed inputs, RE-USE yielded lower word-error-rate point estimates in 11, whereas Denoiser yielded higher estimates in 13. In the paired additive-noise probe, PASE and RE-USE improved recognized-phone agreement, intelligibility, and perceptual-quality point estimates. Denoiser improved recognized-phone agreement and short-time objective intelligibility (STOI) but reduced speaker-embedding similarity. No system was uniformly best across corpora, recognizers, and endpoints. Enhanced rtMRI audio should therefore be treated as a task-specific transformed derivative rather than a universally improved replacement for the original or DSP-processed waveform.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。