arXiv:2409.12477cs.SDcs.AI2024-09中稿 · publication at ICA…被引 4

用音高弯折信息提升小提琴合成的自然感。

ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning

  • 分两阶段:先估计音高弯折,再生成带表现力的频谱
  • 相比无音高弯折建模,声音更真实,听感评分更高
  • 适合关注音乐表现力合成的研究者与创作者

准确建模基频(F0)的自然走势在音乐音频合成中至关重要。然而,在多声部音乐中转录和管理多个F0轮廓仍具挑战性,且目前尚未有研究探索多声部乐器合成中显式建模F0轮廓的方法。本文提出ViolinDiff,一种基于扩散模型的两阶段合成框架。给定小提琴MIDI文件,第一阶段将F0轮廓转化为音高弯折信息进行估计,第二阶段生成包含这些表现细节的梅尔频谱。定量指标与听觉测试结果表明,相比未显式建模音高弯折的模型,所提方法生成的小提琴声音更为真实。音频样例可在线获取:daewoung.github.io/ViolinDiff-Demo。

原文摘要 · Abstract (English)

Modeling the natural contour of fundamental frequency (F0) plays a critical role in music audio synthesis. However, transcribing and managing multiple F0 contours in polyphonic music is challenging, and explicit F0 contour modeling has not yet been explored for polyphonic instrumental synthesis. In this paper, we present ViolinDiff, a two-stage diffusion-based synthesis framework. For a given violin MIDI file, the first stage estimates the F0 contour as pitch bend information, and the second stage generates mel spectrogram incorporating these expressive details. The quantitative metrics and listening test results show that the proposed model generates more realistic violin sounds than the model without explicit pitch bend modeling. Audio samples are available online: daewoung.github.io/ViolinDiff-Demo.

小提琴合成扩散模型音高弯折表现力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。