arXiv:2606.19688cs.SDeess.AS2026-06中稿 · presentation at In…

通过单一超参数实现语音增强的延迟可调,兼顾质量与实时性。

Latency-Configurable Streaming Speech Enhancement via Asymmetric Temporal Padding

论文配图:Latency-Configurable Streaming Speech Enhancement via Asymmetric Temporal Padding
图 1 · 摘自论文原文
  • 采用非对称时间填充和双缓冲机制,灵活配置延迟
  • 12.5~75.0毫秒延迟下PESQ达3.35~3.43,12.5毫秒时超越此前最优
  • 适合需要低延迟且高音质的实时语音处理场景

流式语音增强需在算法延迟与质量间权衡,现有方法多为因果与非因果的二选一。LaCo-SENet通过单个训练时超参数调控两种机制:首先,非对称时间填充重构卷积中的过去与未来上下文;其次,双缓冲流式架构结合状态缓冲(过去)与前瞻缓冲(输入与特征级未来),并引入选择性状态更新,防止未来帧泄漏,确保训练-推理一致性。在VoiceBank+DEMAND数据集上,固定预算(1.37M参数)的骨干模型支持12.5~75.0毫秒延迟,PESQ从3.35提升至3.43。在仅12.5毫秒(完全因果)时,PESQ达3.35,优于此前因果方案(3.27,46.5毫秒)。

原文摘要 · Abstract (English)

Streaming speech enhancement requires balancing algorithmic latency against quality, yet existing approaches largely treat this as a binary causal versus non-causal choice. LaCo-SENet addresses this issue with two mechanisms parameterized by a single training-time hyperparameter. First, asymmetric temporal padding redistributes past and future context in convolutions, enabling systematic latency configuration. Second, dual-buffer streaming combines state buffers for past context with lookahead buffers that supply future context at both the input and feature levels. Selective state updates also prevent future-frame leakage into the streaming state, ensuring training-inference consistency. On VoiceBank+DEMAND, a fixed-budget (1.37M parameters) backbone yields a family of models spanning 12.5-75.0 ms, with PESQ rising from 3.35 to 3.43. At just 12.5 ms (fully causal), a PESQ of 3.35 matches or exceeds the prior causal state-of-the-art (3.27 at 46.5 ms).

语音增强流式处理延迟可控因果性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。