arXiv:2606.18611cs.SDcs.AI2026-06中稿 · Interspeech2026

用四元数结构压缩模型,小体积实现高保真语音增强

QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement

论文配图:QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement
图 1 · 摘自论文原文
  • 用四元数共用权重编码语音幅度与相位,减少参数量
  • 在VoiceBank+DEMAND上仅0.89M参数即达PESQ 3.48分
  • 适合资源受限场景的实时语音增强应用

我们提出一种参数高效的语音增强框架——四元数转换器生成对抗网络(QC-GAN),结合四元数转换器生成器与基于MetricGAN的训练方法。通过哈密顿乘积结构化共享权重,编码语音的幅度与相位信息,在保持二者依赖关系的同时显著减少层参数量。采用度量学习判别器,优化近似感知评估分数以提升听感质量。在VoiceBank+DEMAND数据集上,QC-GAN仅使用0.89M参数即达到3.48的语音感知质量评分(PESQ),性能媲美顶尖模型但规模不足其一半。一个仅35K参数的变体实现PESQ 3.23分,超越传统方法且参数极少。在DNS-Challenge 3数据集上的评估进一步验证了其在真实环境下的泛化能力。

原文摘要 · Abstract (English)

We propose a parameter-efficient speech enhancement framework, Quaternion Conformer GAN (QC-GAN), which combines a Quaternion Conformer generator with MetricGAN-based training. The Hamilton product encodes the magnitude and phase via structured weight sharing, reducing the number of layer parameters while preserving their interdependencies. A metric-learning discriminator was employed to maximize perceptual quality by optimizing the approximate perceptual evaluation scores. On the VoiceBank+DEMAND dataset, QC-GAN achieved a Perceptual Evaluation of Speech Quality (PESQ) score of 3.48 with only 0.89M parameters, delivering a performance comparable to state-of-the-art models at less than half their size. A 35K-parameter variant achieved a PESQ score of 3.23, surpassing conventional methods with significantly fewer parameters. Evaluation on the DNS-Challenge 3 dataset further confirmed generalization to real-world conditions.

语音增强四元数轻量化生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。