arXiv:2608.28127cs.SDeess.AS2026-08中稿 · ISMIR 2026

提出统一框架,区分处理过程与结果音频的表征学习。

Exploring the Design Space of Representation Learning for Audio Transformations

论文配图:Exploring the Design Space of Representation Learning for Audio Transformations
图 1 · 摘自论文原文
  • 设计三重目标统一评估音频变换表征
  • 发现处理嵌入与结果嵌入各擅胜场
  • 适合音频检索与风格迁移研究者

神经音频表征学习已推动多种内容导向应用,但在音频处理任务中仍显不足。现有方法隐含选择仅关注处理过程或保留源内容的结果音频,且模型、数据和评估方式各异,难以判断关键设计影响。本文在统一框架下提出三个目标:处理一致性、描述对齐性与前向预测下的等变性,系统比较所有组合在受控设置下的表现。框架同时生成变换嵌入与处理后音频嵌入,实验发现二者互补:距离相关任务更依赖前者,探针评估任务则偏好后者。结合网络架构与训练流程改进,该表征在检索、探针评估及风格迁移上均超越先前基线。

原文摘要 · Abstract (English)

Neural audio representation learning has enabled a range of content-oriented applications, but the resulting features remain limited for tasks involving audio processing. Furthermore, it is not obvious what processing-aware representations should capture: the processing itself, abstracted away from source content, or the processed audio that retains it. Existing approaches implicitly commit to one or the other and also differ in their models, data, and evaluation, obscuring which design choices drive their behavior. We address both questions within a unified framework of three objectives: processing consistency, description alignment, and equivariance via forward prediction. We compare all combinations of the objectives under a controlled setup and reveal their relative strengths and interactions. Our framework produces both a transformation embedding and a processed-audio embedding, and we find that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter. Combined with improvements in network architecture and training pipeline, our representations outperform prior baselines across retrieval, probe-based evaluation, and style transfer.

音频表征处理感知嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。