arXiv:2602.11896cs.SDeess.AS2026-02

用时频散射生成听感相似但波形不同的音乐幻象

Musical Metamerism with Time--Frequency Scattering

  • 基于时频散射的无预处理音乐幻象生成方法
  • 能保留听觉相似性,不依赖乐谱或节拍分析
  • 适合声音感知、音频生成与神经科学研究者

元色学中的‘色觉融合’描述了两种光谱差异显著却视觉上相似的现象。本文提出‘音乐融合’概念,指两个波形不同但听感相似的音乐片段之间的现象。本技术报告介绍一种从任意音频中生成音乐融合体的方法,基于Kymatio开源库中的联合时频散射(JTFS)。该方法无需手动预处理,如记谱、节拍追踪或声源分离。报告提供JTFS的数学描述及Kymatio部分源码示例,并回顾了相关工作,包括调制功率谱(MPS)、谱时接受场(STRF)和Gabor滤波器组(GBFB)等算法。

原文摘要 · Abstract (English)

The concept of metamerism originates from colorimetry, where it describes a sensation of visual similarity between two colored lights despite significant differences in spectral content. Likewise, we propose to call ``musical metamerism'' the sensation of auditory similarity which is elicited by two music fragments which differ in terms of underlying waveforms. In this technical report, we describe a method to generate musical metamers from any audio recording. Our method is based on joint time--frequency scattering in Kymatio, an open-source software in Python which enables GPU computing and automatic differentiation. The advantage of our method is that it does not require any manual preprocessing, such as transcription, beat tracking, or source separation. We provide a mathematical description of JTFS as well as some excerpts from the Kymatio source code. Lastly, we review the prior work on JTFS and draw connections with closely related algorithms, such as spectrotemporal receptive fields (STRF), modulation power spectra (MPS), and Gabor filterbank (GBFB).

音乐生成时频分析感知模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。