arXiv:2607.11260eess.AS2026-07

用可学习的前端提取语音语义特征,提升低采样率下的重建质量。

Semantic Sampling via Learnable Observation Front Ends

论文配图:Semantic Sampling via Learnable Observation Front Ends
图 1 · 摘自论文原文
  • 通过可学习的语义滤波器和观测矩阵,从波形中生成高信息量观测。
  • 相同采样率下,波形保真度、频谱一致性与听感质量显著优于传统方法。
  • 适合低带宽语音重建、信号压缩等需要高效信息提取的场景。

采样决定了下游重建系统可用的信息形式。传统低率采样直接从原始波形中获取有限维观测,采样规则主要受带宽、稀疏性或固定信号结构影响。但对于语音等声学信号,重建相关的信息往往体现为内容相关的谱-时域结构,而非仅依赖波形样本。本文提出基于可学习观测前段的语义采样方法,使有限维观测由学习得到的信号响应生成,而非直接对波形点进行下采样。该前段包含语义特征滤波器组、受限语义观测矩阵和低率读出模块:滤波器组将输入波形映射到多个声学响应通道,观测矩阵将这些响应融合为少量观测通道,读出模块生成低率有限维样本。随后使用重建网络从观测中恢复信号。在低率语音重建任务上的实验表明,在相同观测预算下,所提语义采样前段生成的观测比固定低率采样和基于预设低率波形的神经恢复方法更具信息量。波形保真度、频谱一致性和感知质量的提升说明,可学习观测前段在同等观测预算下保留了更多有助于声学信号重建的有效信息。

原文摘要 · Abstract (English)

Sampling determines the form of information available to downstream reconstruction systems. Conventional lowrate sampling forms finite-dimensional observations directly from the raw waveform, with the sampling rule mainly guided by bandwidth, sparsity, or fixed signal-level structures. For acoustic signals such as speech, however, reconstruction-relevant information is often expressed through content-related spectral-temporal structures rather than waveform samples alone. This paper proposes semantic sampling via learnable observation front ends, where finite-dimensional observations are generated from learned signal responses instead of directly subsampled waveform points. The proposed front end consists of a semantic feature filterbank, a constrained semantic observation matrix, and a low-rate readout module. The filterbank maps the input waveform into multiple acoustic response channels, the observation matrix combines these responses into a small number of observation channels, and the readout module produces low-rate finite-dimensional samples. A reconstruction network is then used to recover the signal from the resulting observations. Experiments on low-rate speech reconstruction show that, under the same observation budget, the proposed semantic sampling front end provides more informative observations than fixed low-rate sampling and neural restoration methods based on predetermined low-rate waveforms. The improvements in waveform fidelity, spectral consistency, and perceptual quality show that learnable observation front ends preserve more useful information for acoustic signal reconstruction under the same observation budget.

语音重建语义采样可学习前端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。