轻量级神经网络实现高保真双耳音频合成,适合在边缘设备运行。
Lightweight Implicit Neural Network for Binaural Audio Synthesis
- 分两阶段生成:时域扭曲初估 + 隐式修正模块优化
- 参数量比顶尖方法减少72.7%,计算量大幅降低
- 兼顾音质与效率,适合移动端空间音频应用
高保真双耳音频合成对沉浸式听觉体验至关重要,但现有方法计算开销大,限制了其在边缘设备的应用。为此,我们提出轻量级隐式神经网络(Lite-INN),一种新颖的两阶段框架。Lite-INN首先通过时域扭曲生成初始估计,再由隐式双耳校正(IBC)模块进行精修。IBC是一种隐式神经网络,可直接预测振幅和相位修正,从而实现高度紧凑的模型结构。实验表明,Lite-INN在感知质量上与最优基线模型统计上无显著差异,同时显著提升计算效率。相比先前最先进方法(NFS),Lite-INN参数量减少72.7%,所需计算操作(MACs)也大幅降低。这证明该方法有效缓解了音质与计算效率之间的权衡,为高保真边缘设备空间音频应用提供了新方案。
原文摘要 · Abstract (English)
High-fidelity binaural audio synthesis is crucial for immersive listening, but existing methods require extensive computational resources, limiting their edge-device application. To address this, we propose the Lightweight Implicit Neural Network (Lite-INN), a novel two-stage framework. Lite-INN first generates initial estimates using a time-domain warping, which is then refined by an Implicit Binaural Corrector (IBC) module. IBC is an implicit neural network that predicts amplitude and phase corrections directly, resulting in a highly compact model architecture. Experimental results show that Lite-INN achieves statistically comparable perceptual quality to the best-performing baseline model while significantly improving computational efficiency. Compared to the previous state-of-the-art method (NFS), Lite-INN achieves a 72.7% reduction in parameters and requires significantly fewer compute operations (MACs). This demonstrates that our approach effectively addresses the trade-off between synthesis quality and computational efficiency, providing a new solution for high-fidelity edge-device spatial audio applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。