用稀疏事件建模音频,让声音更可解释。
Toward a Sparse and Interpretable Audio Codec
- 将音频表示为稀疏事件及其发生时间,而非传统分块压缩。
- 引入物理规律模拟乐器与空间共振,提升表征可解释性。
- 适合需要理解音频构成的音乐生成与分析场景。
现有主流音频编码器(如 Ogg Vorbis、MP3)及新兴神经编码器(如 Meta Encodec、Descript Audio Codec)均采用分块编码,将音频划分为重叠的固定大小帧进行压缩。此类方法虽能生成高质量音频并支持文本到音频等下游任务,但难以提供直观、可直接解读的表示。本文提出一种概念验证型音频编码器,将音频建模为稀疏事件及其发生时间的集合。通过引入基础物理假设,模拟演奏乐器的起音特性及空间共振效应,旨在获得稀疏、简洁且易于理解的音频表示。
原文摘要 · Abstract (English)
Most widely-used modern audio codecs, such as Ogg Vorbis and MP3, as well as more recent "neural" codecs like Meta's Encodec or the Descript Audio Codec are based on block-coding; audio is divided into overlapping, fixed-size "frames" which are then compressed. While they often yield excellent reproductions and can be used for downstream tasks such as text-to-audio, they do not produce an intuitive, directly-interpretable representation. In this work, we introduce a proof-of-concept audio encoder that represents audio as a sparse set of events and their times-of-occurrence. Rudimentary physics-based assumptions are used to model attack and the physical resonance of both the instrument being played and the room in which a performance occurs, hopefully encouraging a sparse, parsimonious, and easy-to-interpret representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。