用神经网络将4通道全景音频升级为更高精度声音场。
Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
- 直接在波形域用卷积网络从一阶转三阶全景音频
- 定位误差平均低0.6dB,感知质量提升80%
- 适合需要高效又高保真的空间音频应用
全景音频是一种描述声场的空间音频格式。一阶全景音频(FOA)仅含四通道,虽高效但空间精度有限。本文提出一种数据驱动的方案,保留FOA的高效性,同时超越传统渲染器的质量。采用全卷积时域音频神经网络(Conv-TasNet),将输入的FOA转换为更高阶全景音频(HOA)输出。该方法区别于传统的基于物理与心理声学的渲染方式。定量评估显示,预测与真实三阶全景音频之间的平均位置均方误差降低0.6dB;主观评分中,感知质量相较传统方法提升80%。
原文摘要 · Abstract (English)
Ambisonics is a spatial audio format describing a sound field. First-order Ambisonics (FOA) is a popular format comprising only four channels. This limited channel count comes at the expense of spatial accuracy. Ideally one would be able to take the efficiency of a FOA format without its limitations. We have devised a data-driven spatial audio solution that retains the efficiency of the FOA format but achieves quality that surpasses conventional renderers. Utilizing a fully convolutional time-domain audio neural network (Conv-TasNet), we created a solution that takes a FOA input and provides a higher order Ambisonics (HOA) output. This data driven approach is novel when compared to typical physics and psychoacoustic based renderers. Quantitative evaluations showed a 0.6dB average positional mean squared error difference between predicted and actual 3rd order HOA. The median qualitative rating showed an 80% improvement in perceived quality over the traditional rendering approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。