用听觉感知规律提升音乐生成的表达力与连贯性
Expressive Music Data Processing and Generation
- 基于韦伯定律设计听觉感知数据处理,保留音乐细微表达
- 通过概率链规则建模多音符参数间依赖关系,提升音乐连贯性
- 用输出熵筛选稳定可预测乐句,适配音乐美感量化研究
音乐表现力与连贯性在作曲与演奏中至关重要,但常被现代生成模型忽略。本文提出一种基于听觉感知的数据处理技术,源自韦伯定律,反映人类听觉的真实感受,能有效保留音乐的细微表现力。为增强音乐连贯性,我们基于概率链规则,在神经网络中建模音乐数据中多个参数(如音高、时长、力度等)之间的输出依赖关系。实际操作中,将多输出序列模型分解为单输出子模型,并将先前采样结果作为后续子模型的条件,以诱导条件分布。最后,提出一种基于输出熵的初步筛选机制,以熵序列为标准选择可预测且稳定的生成序列。该方法进一步结合信息美学理论,用于量化音乐倾向中的愉悦感与信息增益。
原文摘要 · Abstract (English)
Musical expressivity and coherence are indispensable in music composition and performance, while often neglected in modern AI generative models. In this work, we introduce a listening-based data-processing technique that captures the expressivity in musical performance. This technique derived from Weber's law reflects the human perceptual truth of listening and preserves musical subtlety and expressivity in the training input. To facilitate musical coherence, we model the output interdependencies among multiple arguments in the music data such as pitch, duration, velocity, etc. in the neural networks based on the probabilistic chain rule. In practice, we decompose the multi-output sequential model into single-output submodels and condition previously sampled outputs on the subsequent submodels to induce conditional distributions. Finally, to select eligible sequences from all generations, a tentative measure based on the output entropy was proposed. The entropy sequence is set as a criterion to select predictable and stable generations, which is further studied under the context of informational aesthetic measures to quantify musical pleasure and information gain along the music tendency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。