arXiv:2505.12669cs.SDcs.AI2025-05被引 6

通过推理时对齐提升音乐生成与文本描述的一致性。

Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment

  • 推理阶段引入文本-音频与结构对齐奖励,增强生成音乐与文本匹配度。
  • 在现有模型上不需训练即可提升音乐一致性,客观与主观评分均显著改善。
  • 适合关注文本驱动音乐生成质量的研究者和创作者。

我们提出 Text2midi-InferAlign,一种在推理阶段提升符号化音乐生成质量的新方法。该方法利用文本到音频的对齐信息及音乐结构对齐奖励,在生成过程中强化生成音乐与输入文本描述的一致性。具体地,引入两个目标得分:衡量生成音乐与原始文本在节奏上一致性的文本-音频一致性得分,以及惩罚与调性不符音符的和声一致性得分。通过优化这些基于对齐的目标,在不需额外训练或微调的前提下,使生成的音乐更贴合输入文本,从而提升整体作品的质量与连贯性。我们在现有 Text2midi 模型基础上进行评估,结果表明在客观与主观评价指标上均有显著提升。

原文摘要 · Abstract (English)

We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated music to be consistent with the input caption. Specifically, we introduce two objectives scores: a text-audio consistency score that measures rhythmic alignment between the generated music and the original text caption, and a harmonic consistency score that penalizes generated music containing notes inconsistent with the key. By optimizing these alignment-based objectives during the generation process, our model produces symbolic music that is more closely tied to the input captions, thereby improving the overall quality and coherence of the generated compositions. Our approach can extend any existing autoregressive model without requiring further training or fine-tuning. We evaluate our work on top of Text2midi - an existing text-to-midi generation model, demonstrating significant improvements in both objective and subjective evaluation metrics.

音乐生成文本对齐推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。