arXiv:2510.25714cs.SD2025-10被引 1

用听觉线索分析音频空间质量,无须头模型即可可视化音源分布。

Binaspect -- A Python Library for Binaural Audio Analysis, Visualization & Feature Generation

  • 通过时频域聚类生成可解释的方位图,捕捉声源位置信息。
  • 在比特率、渲染等场景中清晰展现编码失真导致的方位扩散或偏移。
  • 输出结构化特征,适合训练语音质量预测与空间分类模型。

我们提出 Binaspect,一个开源 Python 库,用于双耳音频分析、可视化与特征生成。该工具通过计算改进的双耳时间差与强度差谱图,并将时频单元聚类为稳定的时间-方位直方图表示,从而在不依赖头部模型的前提下,生成可解释的“方位图”。多个活跃声源表现为不同的方位聚类,而编码或渲染失真则体现为分布变宽、弥散或偏移。该方法对音频处理流程中的空间退化具有诊断价值。我们在比特率阶梯、球形麦克风阵列渲染和 VBAP 声源定位任务中验证了其有效性,均能清晰揭示退化现象。此外,所生成的表示可导出为结构化特征,适用于机器学习模型训练,如音频质量预测、空间音频分类等任务。Binaspect 已开源,代码与可复现脚本见 https://github.com/QxLabIreland/Binaspect。

原文摘要 · Abstract (English)

We present Binaspect, an open-source Python library for binaural audio analysis, visualization, and feature generation. Binaspect generates interpretable "azimuth maps" by calculating modified interaural time and level difference spectrograms, and clustering those time-frequency (TF) bins into stable time-azimuth histogram representations. This allows multiple active sources to appear as distinct azimuthal clusters, while degradations manifest as broadened, diffused, or shifted distributions. Crucially, Binaspect operates blindly on audio, requiring no prior knowledge of head models. These visualizations enable researchers and engineers to observe how binaural cues are degraded by codec and renderer design choices, among other downstream processes. We demonstrate the tool on bitrate ladders, ambisonic rendering, and VBAP source positioning, where degradations are clearly revealed. In addition to their diagnostic value, the proposed representations can be exported as structured features suitable for training machine learning models in quality prediction, spatial audio classification, and other binaural tasks. Binaspect is released under an open-source license with full reproducibility scripts at https://github.com/QxLabIreland/Binaspect.

双耳音频空间音频特征提取可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。