arXiv:2607.03806eess.AScs.AI2026-07中稿 · DAFx 2026被引 1

揭示CLAP模型如何编码声音的三大感知属性,发现线性与非线性两种编码方式。

Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings

论文配图:Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
图 1 · 摘自论文原文
  • 用探针模型分析冻结的CLAP音频嵌入,逐层检测声学属性
  • 三种属性在五大数据集上均可准确恢复,谱心率需非线性探针
  • 线性编码具跨数据集一致性,适合音频理解与跨模态研究

音频基础模型被广泛用作通用特征提取器,但其学习表示的内部结构仍不明确。本文通过探针框架分析CLAP音频嵌入,研究三个基本感知维度的编码:混响时间(RT60)、音量(LUFS)和频谱内容(谱心率SC与相对音高RP)。在噪声、语音、单音符音乐及音乐混音等五个数据集上,训练复杂度递增的探针从冻结嵌入中预测各属性。主要发现:所有属性在所考察数据集中均能可靠恢复。整体上出现两种编码模式:RT60、LUFS与RP近似线性编码,而SC需非线性探针。两种模式在八个额外音频基础模型中具有泛化能力,例外是幅度不变架构完全丢弃音量信息。识别出的线性特征方向在RT60和LUFS上跨数据集几何一致,而相对音高方向高度依赖领域。最后,定性展示跨模态一致性:声学描述文本嵌入在几何上对齐于识别出的RT60特征方向。

原文摘要 · Abstract (English)

Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood. In this work, we analyze CLAP audio embeddings through a probing framework, studying the encoding of three fundamental perceptual dimensions: reverberation (RT60), loudness (LUFS), and spectral content, measured via spectral centroid (SC) and relative pitch (RP). Probes of increasing complexity are trained to predict each attribute from frozen embeddings across five datasets spanning noise, speech, monophonic musical notes, and music mixtures. Our primary finding is that all of these attributes are reliably recoverable from the CLAP embedding space across the examined datasets. Within this global picture, two encoding regimes emerge: RT60, LUFS, and RP are approximately linearly encoded, while SC requires non-linear probes. Both regimes generalize across eight additional audio foundation models, with the notable exception that amplitude-invariant architectures discard loudness entirely by construction. The identified linear feature directions are geometrically consistent across datasets for RT60 and LUFS, while highly domain-specific for RP. Finally, we provide a qualitative demonstration of cross-modal consistency, showing that text embeddings of acoustic descriptors align geometrically with the identified RT60 feature direction.

音频嵌入探针分析声学属性跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。