arXiv:2409.06513cs.SDcs.AI2024-09被引 1

用正弦、瞬态、噪声分解法,让合成钢琴音更逼真。

Sines, Transient, Noise Neural Modeling of Piano Notes

  • 分三模块学习谐波、瞬态和噪声,可独立训练。
  • 高频频谱能量预测仍有挑战,但整体频谱分布准确。
  • 适合音乐生成与音频建模研究者使用。

本文提出一种新型钢琴音色模拟方法,通过正弦、瞬态和噪声的分解设计可微分频谱建模合成器,以复现钢琴音符。三个子模块从钢琴录音中学习并生成对应的谐波、瞬态和噪声信号。将建模任务拆分为三个独立可训练模型,降低复杂度。准谐波成分采用基于物理公式的可微正弦模型生成,参数由音频自动估计;噪声子模块使用可学习的时变滤波器,瞬态则由深度卷积网络生成。针对三音和弦,利用基于卷积网络的模型模拟不同琴键间的耦合效应。结果表明,模型能匹配目标音符的泛音分布,但在高频段能量预测仍具挑战;瞬态与噪声成分的频谱能量分布总体准确。尽管计算与内存效率更高,听觉测试显示其在音符起始阶段的建模仍有不足,但对单音及三音和弦的整体感知保真度良好。

原文摘要 · Abstract (English)

This paper introduces a novel method for emulating piano sounds. We propose to exploit the sines, transient, and noise decomposition to design a differentiable spectral modeling synthesizer replicating piano notes. Three sub-modules learn these components from piano recordings and generate the corresponding harmonic, transient, and noise signals. Splitting the emulation into three independently trainable models reduces the modeling tasks' complexity. The quasi-harmonic content is produced using a differentiable sinusoidal model guided by physics-derived formulas, whose parameters are automatically estimated from audio recordings. The noise sub-module uses a learnable time-varying filter, and the transients are generated using a deep convolutional network. From singular notes, we emulate the coupling between different keys in trichords with a convolutional-based network. Results show the model matches the partial distribution of the target while predicting the energy in the higher part of the spectrum presents more challenges. The energy distribution in the spectra of the transient and noise components is accurate overall. While the model is more computationally and memory efficient, perceptual tests reveal limitations in accurately modeling the attack phase of notes. Despite this, it generally achieves perceptual accuracy in emulating single notes and trichords.

音频建模可微分合成钢琴音色

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。