提出新型语音修复框架,支持输入输出采样率解耦。
Query-Based Asymmetric Modeling with Decoupled Input-Output Rates for Speech Restoration
- 用查询机制分离输入分析与输出合成,实现非对称建模。
- 在多采样率下统一处理降噪、混响消除等任务,性能均衡。
- 无需冗余重采样,适合实际语音修复场景应用。
语音修复旨在从受噪声、混响、带宽限制等失真影响的录音中恢复干净语音,其输入与输出采样率可能不同。现有方法通常假设输入输出采样率一致,并采用冗余重采样,限制了原生多速率处理能力。本文将此问题形式化为扩展的采样频率无关(xSFI)设置,提出TF-Restormer——一种基于查询的xSFI建模框架。该模型仅编码观测到的输入频带,通过带分区交叉注意力的扩展查询合成未观测的高频部分,形成不对称编码器-解码器结构,将容量分配给分析端,保持合成端轻量。训练采用感知损失、缩放对数谱损失及SFI-STFT判别器的对抗监督,使模型在单个统一架构下实现保真度与感知质量的平衡,在多种采样率下,于降噪、去混响、带宽扩展及综合失真基准上均表现优异,且无需冗余重采样。
原文摘要 · Abstract (English)
Speech restoration aims to recover clean speech from degraded recordings affected by noise, reverberation, bandwidth reduction, or other distortions, where input and output sampling rates may differ. Existing approaches typically assume matched input-output rates and apply redundant resampling, limiting native multi-rate processing. We formulate this gap as the extended sampling-frequency-independent (xSFI) setting, where a model must operate under decoupled input-output rates, and propose TF-Restormer, a query-based xSFI modeling framework. The model encodes only the observed input band and synthesizes the unobserved high-frequency band through extension queries with band-partitioned cross-attention, yielding an asymmetric encoder-decoder that allocates capacity to analysis while keeping synthesis lightweight. Trained with a perceptual loss, a scaled log-spectral loss, and adversarial supervision via an SFI-STFT discriminator, TF-Restormer attains balanced fidelity-perceptual quality as a single unified model, without redundant resampling across denoising, dereverberation, bandwidth extension, and combined distortion benchmarks under multiple sampling rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。