arXiv:2606.07207cs.SDcs.LG2026-06

用输出能量熵自动调节训练梯度,提升音乐生成多样性。

Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development

  • 基于DiT输出空间能量分布的熵设计无参数梯度权重
  • 在MusicCaps上实现更强主题发展与更丰富音色差异
  • 无需调参,自动生成动态训练数据课程

置信度加权通常被避免用于生成模型,因其在模型自信出错时会加速错误传播,但在监督扩散训练中此直觉失效。本文提出Eisbach对数障碍项,一种源自DiT输出空间能量分布熵的无参数权重:高熵抑制梯度,低熵保留梯度。将其应用于Stable Audio 3 Medium在MusicCaps上的LoRA微调,意外获得更强的主题发展、更清晰的声学区分及更高的纹理多样性,与模式崩溃相反。这是因为监督扩散中梯度方向由真实标签锁定,置信度仅影响步长;而时间熵会降低平坦样本权重,保留高对比样本。结果是仅通过前向传播即可生成在线自指数据课程,并分析了噪声水平动态,给出可验证预测。

原文摘要 · Abstract (English)

Confidence-based loss weighting is usually avoided in generative models because it accelerates errors when the model is confidently wrong, but this intuition breaks down in supervised diffusion training. We introduce the Eisbach log-barrier, a parameter-free weight derived from the entropy of the DiT output's spatial energy distribution: high entropy damps the gradient, while low entropy preserves it. Applied to LoRA fine-tuning of Stable Audio 3 Medium on MusicCaps, it unexpectedly yields stronger thematic development, clearer acoustic differentiation, and higher textural diversity than unweighted training, the opposite of mode collapse. This works because in supervised diffusion the gradient direction is locked to ground truth, so confidence only scales the step size, and because temporal entropy downweights flat samples while preserving high-contrast ones. The result is an online, self-referential data curriculum that emerges purely from the forward pass, with analyzed noise-level dynamics and testable predictions.

扩散模型音乐生成自适应训练熵正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。