arXiv:2603.13685cs.SD2026-03中稿 · ICASSP 2026

首个评估音频表征组合性的基准,检验模型能否拆解与重组声音元素。

Evaluating Compositional Structure in Audio Representations

  • 设计两类任务:加性变换一致性测试与属性级重建测试
  • 基于可控声学属性的合成数据集,支持大规模评估
  • 适合研究音频语义理解与生成的学者使用

我们提出一个评估音频表征组合性的基准。音频组合性指将声音场景分解为组成部分与属性,并系统化地组合它们。尽管这在听觉感知中至关重要,但当前评估协议大多忽略这一特性。我们的框架通过A-COAT任务测试加性变换下的表现一致性,通过A-TRE任务探查从属性级原始元素中重建的能力。两项任务均基于包含受控声学属性变化的大规模合成数据集,首次实现了对音频嵌入组合结构的系统性评估。

原文摘要 · Abstract (English)

We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attributes, and combining them systematically. While central to auditory perception, this property is largely absent from current evaluation protocols. Our framework adapts ideas from vision and language to audio through two tasks: A-COAT, which tests consistency under additive transformations, and A-TRE, which probes reconstructibility from attribute-level primitives. Both tasks are supported by large synthetic datasets with controlled variation in acoustic attributes, providing the first benchmark of compositional structure in audio embeddings.

音频表征组合性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。