arXiv:2505.18984cs.SDeess.AS2025-05被引 2

用多种采样策略提升音频表示通用性,改善音高检测等任务表现。

Self-supervised learning method using multiple sampling strategies for general-purpose audio representation

  • 融合片段、帧级和任务特定三种采样策略构建对比损失。
  • 在冻结权重下,音高检测性能提升3.6%,片段分类提升25%。
  • 适合需要强泛化能力的音频理解任务,如事件检测与音高识别。

我们提出一种基于多采样策略的自监督学习方法,以获得通用音频表征。除广泛使用的片段级采样外,还引入帧级采样和任务特定采样两种新策略,从不同角度构建对比损失并学习表征。该方法在Audioset子集上预训练,并在下游任务中使用冻结权重。结果表明,该方法在片段分类、声音事件检测和音高检测任务上的性能分别提升了25%、20%和3.6%。

原文摘要 · Abstract (English)

We propose a self-supervised learning method using multiple sampling strategies to obtain general-purpose audio representation. Multiple sampling strategies are used in the proposed method to construct contrastive losses from different perspectives and learn representations based on them. In this study, in addition to the widely used clip-level sampling strategy, we introduce two new strategies, a frame-level strategy and a task-specific strategy. The proposed multiple strategies improve the performance of frame-level classification and other tasks like pitch detection, which are not the focus of the conventional single clip-level sampling strategy. We pre-trained the method on a subset of Audioset and applied it to a downstream task with frozen weights. The proposed method improved clip classification, sound event detection, and pitch detection performance by 25%, 20%, and 3.6%.

自监督学习音频表征多采样策略音高检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。