arXiv:2603.02794cs.SDcs.AI2026-03

用可解释的滤波器实时降噪,让助听设备更智能可控。

An Interpretable, Controllable Time-Varying IIR Denoiser for On-Device Assistive Hearing

  • 用神经控制器动态调节35个二阶IIR滤波器,实现可解释的实时降噪。
  • 仅24k参数、10.7ms延迟,性能接近230万参数大模型。
  • 支持运行时调节降噪强度,无需重训练,适合助听设备部署。

我们提出TVF(Time-Varying Filtering),一种用于实时、设备端辅助听力的可解释、低延迟语音增强模型。一个轻量级神经控制器实时预测35个二阶IIR滤波器(biquads)的系数,使模型能跟踪非平稳噪声,同时保持完全可解释的处理链:每个频谱修改都是显式可调的均衡曲线,而非黑箱变换。由于信号处理由滤波器级联承担,控制器可做到极小规模,仅需24k参数,算法延迟为10.7ms,符合助听器预算,且全程在设备端运行,音频不外泄。我们还将抑制与保留的权衡作为显式控制:可在训练中通过损失权重设定,推理时通过混合原始噪声输入与去噪输出调整,无需重训练。在助听器指标(HASPI/HASQI)上,24k模型性能比DFNet3(230万参数,几乎大两个数量级)低约0.02,但乘累加操作减少约29倍;尽管更大黑箱模型在参考指标如PESQ上仍领先。本文将TVF作为紧凑、可解释、可控去噪器在设备端辅助听力应用的可行性验证。

原文摘要 · Abstract (English)

We present TVF (Time-Varying Filtering), an interpretable, low-latency speech enhancement model for real-time, on-device assistive hearing. A lightweight neural controller predicts, in real time, the coefficients of a differentiable cascade of 35 second-order IIR filters (biquads), so the model tracks non-stationary noise while keeping a fully interpretable processing chain: every spectral modification is an explicit, adjustable equalizer curve rather than an opaque `black-box' transform. Because the biquad cascade carries the signal processing, the controller can be made very small, driving the cascade with only 24k parameters at a 10.7ms algorithmic latency, within hearing-aid budgets, and running entirely on-device so that audio never leaves the device. We also expose the suppression-versus-preservation trade-off as an explicit control: it can be set during training through the loss weighting, and adjusted at inference, with no retraining, by mixing the noisy input with the denoised output. On hearing-aid metrics (HASPI/HASQI) the 24k model stays within about 0.02 of DFNet3 (2.3M parameters, almost two orders of magnitude larger) while using about 29X fewer multiply-accumulates, although larger black-box models still lead on reference metrics such as PESQ. We present TVF as a proof of concept for a compact, interpretable, and controllable denoiser for on-device assistive hearing.

语音增强助听设备可解释性低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。