用音频特征控制扩散模型生成同步音乐视频。
Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation
- 结合风格描述与音频能量向量,用扩散模型生成视觉图像。
- 新指标AVS显示该方法显著提升音画同步性。
- 适合音乐可视化、演出和公共空间视听体验应用。
本研究提出一种基于扩散模型的音乐视觉化生成方法,结合音频输入与用户选定的艺术作品。流程分为图像生成与视频创建两阶段:首先进行音乐标题生成与流派分类,随后检索艺术风格描述;扩散模型根据用户输入图像与提取的风格描述生成图像。视频生成阶段采用相同扩散模型进行帧间插值,由谐波与打击乐特征提取的音频能量向量控制。实验表明该方法在多种音乐流派中表现良好,引入新评价指标音频-视觉同步性(AVS),对比结果显示使用音频能量向量生成的视频在AVS上显著高于线性插值。该方法适用于独立音乐视频创作、影视制作、现场演出及公共空间音视频体验增强。
原文摘要 · Abstract (English)
This study presents a novel method for generating music visualisers using diffusion models, combining audio input with user-selected artwork. The process involves two main stages: image generation and video creation. First, music captioning and genre classification are performed, followed by the retrieval of artistic style descriptions. A diffusion model then generates images based on the user's input image and the derived artistic style descriptions. The video generation stage utilises the same diffusion model to interpolate frames, controlled by audio energy vectors derived from key musical features of harmonics and percussives. The method demonstrates promising results across various genres, and a new metric, Audio-Visual Synchrony (AVS), is introduced to quantitatively evaluate the synchronisation between visual and audio elements. Comparative analysis shows significantly higher AVS values for videos generated using the proposed method with audio energy vectors, compared to linear interpolation. This approach has potential applications in diverse fields, including independent music video creation, film production, live music events, and enhancing audio-visual experiences in public spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。