arXiv:2512.10375cs.SDcs.AI2025-12

用神经网络实现可灵活调整的个人声区,一次训练搞定多种听音场景。

Neural personal sound zones with flexible bright zone control

  • 用3D卷积神经网络直接生成声区预滤波器,输入虚拟声场输出控制参数。
  • 仅需一次训练即可适配不同播放目标和任意控制点布局,提升实用性。
  • 从稀疏采样点中学习全局空间信息,减少对密集测量的依赖。

个人声区(PSZ)系统通过单一扬声器阵列在相同空间内为不同位置的听众创建独立的虚拟声场,是虚拟现实应用中的关键技术。传统方法需在与录制房间脉冲响应(RIRs)相同的固定接收阵列上测量重建目标,导致实际部署成本高、操作不便。本文提出一种基于3D卷积神经网络的PSZ重建方法,以虚拟目标声场为输入,输出对应声区预滤波器,支持灵活控制麦克风网格和可变再现目标。实验表明,该方法仅需一次训练即可处理多种再现目标和灵活的控制点布局;同时具备从分布于各声区内稀疏采样点中学习全局空间信息的能力,显著降低对密集测量的依赖。

原文摘要 · Abstract (English)

Personal sound zone (PSZ) reproduction system, which attempts to create distinct virtual acoustic scenes for different listeners at their respective positions within the same spatial area using one loudspeaker array, is a fundamental technology in the application of virtual reality. For practical applications, the reconstruction targets must be measured on the same fixed receiver array used to record the local room impulse responses (RIRs) from the loudspeaker array to the control points in each PSZ, which makes the system inconvenient and costly for real-world use. In this paper, a 3D convolutional neural network (CNN) designed for PSZ reproduction with flexible control microphone grid and alternative reproduction target is presented, utilizing the virtual target scene as inputs and the PSZ pre-filters as output. Experimental results of the proposed method are compared with the traditional method, demonstrating that the proposed method is able to handle varied reproduction targets on flexible control point grid using only one training session. Furthermore, the proposed method also demonstrates the capability to learn global spatial information from sparse sampling points distributed in PSZs.

声区控制神经网络音频渲染空间音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。