arXiv:2503.02585cs.SDcs.CV2025-03中稿 · the 34th ACM Inter…被引 2

KAN用可学习激活函数实现高保真音频表征,效果优于现有方法。

As Good as It KAN Get: High-Fidelity Audio Representation

  • 用可学习激活函数的KAN架构,提升音频隐式表征质量。
  • 1.5秒音频上达到1.29的最低频谱距离和3.57的最高语音感知评分。
  • FewSound增强参数更新,相比SOTA提升超60%的语音清晰度。

隐式神经表示(INR)在多媒体数据编码中表现出高效性,但在音频信号中的应用仍有限。本文提出基于可学习激活函数的柯尔莫戈罗夫-阿诺德网络(KAN),作为高效的音频表征INR模型。KAN在感知性能上优于以往INR,在1.5秒音频上取得1.29的最低对数谱距离和3.57的最高语音感知质量评分。为进一步拓展其应用,提出基于超网络的FewSound架构,优化INR参数更新。FewSound在均方误差上较SOTA HyperSound提升33.3%,信噪比提升60.87%。结果表明KAN具备强鲁棒性与可扩展性,适合集成至各类超网络框架。源码见https://github.com/gmum/fewsound.git。

原文摘要 · Abstract (English)

Implicit neural representations (INR) have gained prominence for efficiently encoding multimedia data, yet their applications in audio signals remain limited. This study introduces the Kolmogorov-Arnold Network (KAN), a novel architecture using learnable activation functions, as an effective INR model for audio representation. KAN demonstrates superior perceptual performance over previous INRs, achieving the lowest Log-SpectralDistance of 1.29 and the highest Perceptual Evaluation of Speech Quality of 3.57 for 1.5 s audio. To extend KAN's utility, we propose FewSound, a hypernetwork-based architecture that enhances INR parameter updates. FewSound outperforms the state-of-the-art HyperSound, with a 33.3% improvement in MSE and 60.87% in SI-SNR. These results show KAN as a robust and adaptable audio representation with the potential for scalability and integration into various hypernetwork frameworks. The source code can be accessed at https://github.com/gmum/fewsound.git.

音频表征隐式表示超网络可学习激活

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。