arXiv:2607.01669cs.SD2026-07中稿 · ICME 2026 Grand Ch…被引 1

通过文本嵌入聚类提升小数据下的文本音乐生成效果

UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation

  • 用文本或音频嵌入对训练数据聚类,同质样本放入同一批次以减少梯度干扰
  • 基于文本嵌入的聚类在客观指标上优于音频嵌入,中等簇数在指标上最佳
  • 更多簇数生成的音乐结构更连贯,适合听感评估

本文研究在低数据量和小规模模型设置下,训练时批量采样策略对文本到音频音乐生成的影响。采用文本嵌入或音频嵌入对训练数据进行聚类,将特征相似的样本分组至同一小批量,以缓解梯度干扰。分析了模态选择与聚类粒度的影响。结果表明,基于文本嵌入的聚类在客观评价指标上表现优于基于音频嵌入;不同聚类粒度导致不同表现:中等数量的簇在客观指标上最优,而更多簇则在听觉测试中展现出更连贯的音乐结构。

原文摘要 · Abstract (English)

This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describes our approach and findings for the ICME 2026 Grand Challenge on Academic Text-to-Music Generation. Training data are clustered using either text embeddings or audio embeddings, and samples with similar characteristics are grouped within the same mini-batch to mitigate gradient interference. The effects of modality and cluster granularity on clustering are analyzed. Results show that clustering based on text embeddings achieves better performance on objective evaluation metrics than clustering based on audio embeddings. In addition, different cluster granularity leads to different behaviors across evaluation criteria: a moderate number of clusters performs best on objective metrics, while a larger number of clusters tends to exhibit music with more coherent structure in listening tests.

文本生成音乐小样本学习聚类采样音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。