统一处理任意空间音频采集与播放格式的转码框架。
Generalised Transcoding Framework for Arbitrary Spatial Audio Capture and Playback Formats

- 基于时频依赖的空间元数据,建模声源与环境声特性。
- 支持低阶和几何受限麦克风阵列,提升听感质量。
- 适合需要灵活适配不同设备的音频系统开发者。
本文提出一种统一框架,用于对以全向声学信号或原始麦克风阵列信号捕获的空间音频场景进行参数化分析与重放。该方法估计时频依赖的空间元数据,表征可变数量的主要声源分量及具有独立方向功率分布的环境声分量,其参数拟合所观测到的信号空间协方差。利用这些元数据构建目标重放格式的空间协方差,并推导出最优混合矩阵,实现场景向目标播放系统的转码。该方法还支持采集与重放装置的独立旋转。在使用全向、球面及头戴阵列模拟场景的主观听觉测试中,对比了本方法与现有先进参数渲染器的实时实现。结果表明,在多种内容和接收配置下,该框架表现出显著感知优势,尤其在低阶及几何受限麦克风阵列中表现更优。
原文摘要 · Abstract (English)
This article introduces a unified framework for the parametric analysis and reproduction of spatial sound scenes captured either as Ambisonic signals or as raw microphone array signals. The proposed method estimates time-frequency-dependent spatial metadata that characterises a variable number of primary source components and an ambience component with its own angular power distribution, whose parameters fit the observed spatial covariances of the captured signals. This metadata is used to construct spatial covariances of the target playback formats, which are then used to derive optimal mixing matrices for transcoding the scene for playback over the target reproduction system. The method additionally handles independent rotations of both capture and playback setups. Real-time implementations of the method and other existing state-of-the-art parametric renderers are compared in a listening test using simulated scenes from Ambisonic, spherical, and head-worn arrays. The results highlight perceptual benefits of the proposed framework across a diverse range of content and receiver configurations, particularly for lower-order and geometrically constrained microphone arrays.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。