arXiv:2601.18322eess.AScs.SD2026-01被引 2

用神经网络修正线性编码器,提升智能眼镜的沉浸式音频质量

Residual Learning for Neural Ambisonics Encoders

  • 以线性编码器为基础,用神经网络做残差修正
  • 在真实数据上,两种神经模型均显著优于传统方法
  • 适合做可穿戴设备空间音频编码的研究与开发者

智能眼镜等可穿戴设备对高质量空间音频捕获提出需求,紧凑的头戴麦克风阵列难以实现精准编码。传统线性编码器虽鲁棒但易放大低频噪声并产生高频空间混叠。神经网络方法性能更优,但常依赖理想麦克风假设,现实表现不稳定。本文提出一种残差学习框架,用神经网络对线性编码器进行修正。基于智能眼镜实测阵列传输函数,对比了文献中的UNet模型和新型循环注意力模型。结果表明:仅当嵌入残差框架时,两类神经编码器才在所有测试指标上持续超越线性基线,域内数据表现显著提升,域外数据也有适度增益。但相干性分析显示,各类神经模型在高频方向保真度方面仍存局限。

原文摘要 · Abstract (English)

Emerging wearable devices such as smartglasses and extended reality headsets demand high-quality spatial audio capture from compact, head-worn microphone arrays. Ambisonics provides a device-agnostic spatial audio representation by mapping array signals to spherical harmonic (SH) coefficients. In practice, however, accurate encoding remains challenging. While traditional linear encoders are signal-independent and robust, they amplify low-frequency noise and suffer from high-frequency spatial aliasing. On the other hand, neural network approaches can outperform linear encoders but they often assume idealized microphones and may perform inconsistently in real-world scenarios. To leverage their complementary strengths, we introduce a residual-learning framework that refines a linear encoder with corrections from a neural network. Using measured array transfer functions from smartglasses, we compare a UNet-based encoder from the literature with a new recurrent attention model. Our analysis reveals that both neural encoders only consistently outperform the linear baseline when integrated within the residual learning framework. In the residual configuration, both neural models achieve consistent and significant improvements across all tested metrics for in-domain data and moderate gains for out-of-domain data. Yet, coherence analysis indicates that all neural encoder configurations continue to struggle with directionally accurate high-frequency encoding.

空间音频神经编码残差学习可穿戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。