Zimtohrli用生物听觉模型提升音频相似度评估效率与准确性
Zimtohrli: An Efficient Psychoacoustic Audio Similarity Metric
- 基于耳蜗频率分辨率和鼓膜非线性响应构建听觉感知前端
- 在公开数据集上超越ViSQOL,接近商业级POLQA性能
- 适合音频编解码、生成音频评估等需要快速准确的场景
本文提出Zimtohrli,一种新型全参考音频相似度度量方法,旨在实现高效且符合人类听觉感知的质量评估。在深度学习模型计算开销大、传统标准专有封闭的背景下,Zimtohrli通过128通道的伽马音调滤波器组模拟耳蜗频率分辨率,并结合独特的非线性谐振器模型,模仿人耳鼓膜对声学刺激的响应。通过改进的动态时间规整(DTW)与神经图相似度指数(NSIM)算法,比较经感知映射的频谱图,引入新非线性以更贴近人类判断。Zimtohrli在性能上优于开源基准ViSQOL,并显著缩小与最新商用标准POLQA的差距。其在感知相关性与计算效率间取得良好平衡,可作为现代音频工程应用(如编解码器开发、生成音频系统评估)的有力替代方案。
原文摘要 · Abstract (English)
This paper introduces Zimtohrli, a novel, full-reference audio similarity metric designed for efficient and perceptually accurate quality assessment. In an era dominated by computationally intensive deep learning models and proprietary legacy standards, there is a pressing need for an interpretable, psychoacoustically-grounded metric that balances performance with practicality. Zimtohrli addresses this gap by combining a 128-bin gammatone filterbank front-end, which models the frequency resolution of the cochlea, with a unique non-linear resonator model that mimics the human eardrum's response to acoustic stimuli. Similarity is computed by comparing perceptually-mapped spectrograms using modified Dynamic Time Warping (DTW) and Neurogram Similarity Index Measure (NSIM) algorithms, which incorporate novel non-linearities to better align with human judgment. Zimtohrli achieves superior performance to the baseline open-source ViSQOL metric, and significantly narrows the performance gap with the latest commercial POLQA metric. It offers a compelling balance of perceptual relevance and computational efficiency, positioning it as a strong alternative for modern audio engineering applications, from codec development to the evaluation of generative audio systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。