arXiv:2601.03612cs.LGcs.SD2026-01

用数学方法让音乐生成更连贯,减少48%参数量

Mathematical Foundations of Polyphonic Music Generation via Structural Inductive Bias

  • 将音高与手部动作解耦建模,基于实证独立性设计结构化嵌入
  • 参数量减少48.3%,验证损失降低9.47%,泛化能力提升28.09%
  • 适合关注音乐生成可解释性与高效建模的研究者

本文针对人工智能音乐生成中的‘缺失中间层’问题——即难以生成连贯的乐句结构——以贝多芬钢琴奏鸣曲为案例,提出智能嵌入(Smart Embedding)架构。该架构基于音高与手部属性在经验上独立的发现(NMI=0.167),采用因子化表示,在保持性能的同时使嵌入参数减少48.3%,并使验证损失下降9.47%。理论上,通过信息论、Rademacher复杂度分析(泛化界收紧28.09%)及范畴论解释建立形式化保证。支持性证据包括奇异值分解分析和盲听专家评测(N=53)。本工作融合架构创新与数学严谨性,为复杂序列数据生成提供可解释、高效且稳定的原理框架。

原文摘要 · Abstract (English)

This monograph addresses the "Missing Middle" problem in AI music generation - the challenge of producing coherent, phrase-level musical structure. Using Beethoven's piano sonatas as a case study, I introduce the Smart Embedding architecture, a factorized representation grounded in the empirically verified independence of pitch and hand attributes (NMI=0.167). The architecture achieves a 48.3% reduction in embedding parameters while improving validation loss by 9.47%. Theoretically, I establish formal guarantees through information theory, Rademacher complexity analysis (yielding a 28.09% tighter generalization bound), and category-theoretic interpretation. These results are further supported by Singular Value Decomposition analysis and a blind expert listening study (N=53). Collectively, this work presents a dual contribution that combines architectural innovation with mathematical rigor, offering a principled framework for building more efficient, stable, and interpretable generative models for complex sequential data.

音乐生成结构先验嵌入压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。