仅凭音频生成逼真小提琴演奏动作,无需乐谱或MIDI。
SyncViolinist: Music-Oriented Violin Motion Generation Based on Bowing and Fingering
- 分两阶段端到端框架,从音频提取弓法指法信息
- 生成动作在时间粒度与演奏特性上更精准
- 专业演奏者评测验证效果,适合音乐动画与虚拟演出
自动生成逼真的音乐表演动作可显著提升数字媒体制作水平,但捕捉小提琴演奏所需的全身、手部及手指复杂运动仍具挑战。现有方法常因音频与动作间映射复杂,需额外输入如乐谱或MIDI数据。本文提出SyncViolinist,一种仅依赖音频输入的多阶段端到端框架,通过弓法/指法模块提取音频中的演奏细节,由动作生成模块生成精确协调的全身动作,反映小提琴演奏的时间粒度与表现特征。实验表明,该方法在未见过的小提琴演奏音频上取得显著优于现有技术的定性与定量结果,专业演奏者主观评估进一步验证其有效性。代码与数据集已公开于https://github.com/Kakanat/SyncViolinist。
原文摘要 · Abstract (English)
Automatically generating realistic musical performance motion can greatly enhance digital media production, often involving collaboration between professionals and musicians. However, capturing the intricate body, hand, and finger movements required for accurate musical performances is challenging. Existing methods often fall short due to the complex mapping between audio and motion, typically requiring additional inputs like scores or MIDI data. In this work, we present SyncViolinist, a multi-stage end-to-end framework that generates synchronized violin performance motion solely from audio input. Our method overcomes the challenge of capturing both global and fine-grained performance features through two key modules: a bowing/fingering module and a motion generation module. The bowing/fingering module extracts detailed playing information from the audio, which the motion generation module uses to create precise, coordinated body motions reflecting the temporal granularity and nature of the violin performance. We demonstrate the effectiveness of SyncViolinist with significantly improved qualitative and quantitative results from unseen violin performance audio, outperforming state-of-the-art methods. Extensive subjective evaluations involving professional violinists further validate our approach. The code and dataset are available at https://github.com/Kakanat/SyncViolinist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。