用双层束搜索实现音乐张力的精准控制,让生成旋律更符合预期情绪。
Explicit Tonal Tension Conditioning via Dual-Level Beam Search for Symbolic Music Generation
- 引入音程向量分析的张力模型,结合双层束搜索调节音乐张力。
- 在多个测试中成功匹配目标张力曲线,主观评价认可度高。
- 同一张力条件下可生成多种不同风格的音乐变体,适合创作辅助。
当前最先进的符号化音乐生成模型虽已达到高质量输出,但对调性张力等作曲特征的显式控制仍具挑战。本文提出一种新方法,将基于音程向量分析的计算张力模型融入Transformer框架,并在推理阶段采用双层束搜索策略。在标记层,通过模型概率与多样性度量对候选序列重新排序,以保持整体质量;在小节层,则基于张力指标进行重排序,确保生成音乐符合预设张力曲线。客观评估显示,该方法能有效调控调性张力,主观听感测试也证实生成结果与目标张力一致。实验还表明,同一张力条件下可生成多个风格各异的音乐版本,证明该方法为引导人工智能作曲提供了强大而直观的工具。
原文摘要 · Abstract (English)
State-of-the-art symbolic music generation models have recently achieved remarkable output quality, yet explicit control over compositional features, such as tonal tension, remains challenging. We propose a novel approach that integrates a computational tonal tension model, based on tonal interval vector analysis, into a Transformer framework. Our method employs a two-level beam search strategy during inference. At the token level, generated candidates are re-ranked using model probability and diversity metrics to maintain overall quality. At the bar level, a tension-based re-ranking is applied to ensure that the generated music aligns with a desired tension curve. Objective evaluations indicate that our approach effectively modulates tonal tension, and subjective listening tests confirm that the system produces outputs that align with the target tension. These results demonstrate that explicit tension conditioning through a dual-level beam search provides a powerful and intuitive tool to guide AI-generated music. Furthermore, our experiments demonstrate that our method can generate multiple distinct musical interpretations under the same tension condition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。