arXiv:2409.11498cs.SDcs.AI2024-09被引 6

通过增强文本多样性提升音乐-文本表示学习效果

Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning

  • 提出增广视图丢弃与文本替换技术,提升训练文本多样性
  • 在有限资源下,数据清洗比模型选择更重要
  • 无需额外数据或计算,通用性强,适合实际部署

音频-文本对比模型已成为音乐表征学习的有力方法。尽管其表现出色,但关键设计选择对音乐-文本表征质量的影响仍不清楚。本文在数据与计算资源受限的条件下,系统考察了三大因素:基础编码器选择、训练数据的筛选程度以及文本增广策略。实验发现,在资源受限场景下,数据清洗是决定性能的最关键因素。基于此洞察,我们提出两种新方法:增广视图丢弃(Augmented View Dropout)和TextSwap,有效提升训练时所见文本的多样性和描述性。实验表明,这些方法在不同预训练方案、模型架构及下游数据分布下均能提升性能,且不增加计算开销或需要额外训练数据。

原文摘要 · Abstract (English)

Audio-text contrastive models have become a powerful approach in music representation learning. Despite their empirical success, however, little is known about the influence of key design choices on the quality of music-text representations learnt through this framework. In this work, we expose these design choices within the constraints of limited data and computation budgets, and establish a more solid understanding of their impact grounded in empirical observations along three axes: the choice of base encoders, the level of curation in training data, and the use of text augmentation. We find that data curation is the single most important factor for music-text contrastive training in resource-constrained scenarios. Motivated by this insight, we introduce two novel techniques, Augmented View Dropout and TextSwap, which increase the diversity and descriptiveness of text inputs seen in training. Through our experiments we demonstrate that these are effective at boosting performance across different pre-training regimes, model architectures, and downstream data distributions, without incurring higher computational costs or requiring additional training data.

音乐表征对比学习文本增广高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。