arXiv:2601.12180cs.HCcs.MM2026-01中稿 · CHI 2026被引 1

用上下文缩略图+自然语言编辑,让视频配乐创作更直观高效。

VidTune: Creating Video Soundtracks with Generative Music and Contextual Thumbnails

  • 根据视频内容生成带视觉提示的音乐选项,支持快速对比。
  • 通过颜色亮度等视觉特征映射音乐情绪与节奏,提升感知一致性。
  • 支持自然语言修改并实时生成新曲目,适合创意创作者使用。

音乐塑造视频氛围,但创作者常难以找到契合视频情绪与叙事的配乐。尽管文本生成音乐模型已出现,我们的前期研究(N=8)发现创作者在构建多样化提示、快速审阅比较曲目及理解音乐对视频影响方面存在困难。本文提出VidTune系统,通过创作者输入的提示生成多样音乐,并生成基于上下文的缩略图以实现快速评估。系统提取视频关键主体作为视觉基础,将每首曲目的情感值(valence)和能量(energy)映射为色彩与明暗等视觉线索,并可视化主要流派与乐器。创作者可通过自然语言修改曲目,VidTune将其扩展为新生成版本。在受控用户研究(N=12)与探索性案例研究(N=6)中,参与者普遍认为VidTune有助于高效审阅与比较音乐,且过程充满趣味与创造性。

原文摘要 · Abstract (English)

Music shapes the tone of videos, yet creators often struggle to find soundtracks that match their video's mood and narrative. Recent text-to-music models let creators generate music from text prompts, but our formative study (N=8) shows creators struggle to construct diverse prompts, quickly review and compare tracks, and understand their impact on the video. We present VidTune, a system that supports soundtrack creation by generating diverse music options from a creator's prompt and producing contextual thumbnails for rapid review. VidTune extracts representative video subjects to ground thumbnails in context, maps each track's valence and energy onto visual cues like color and brightness, and depicts prominent genres and instruments. Creators can refine tracks through natural language edits, which VidTune expands into new generations. In a controlled user study (N=12) and an exploratory case study (N=6), participants found VidTune helpful for efficiently reviewing and comparing music options and described the process as playful and enriching.

视频配乐生成音乐人机交互视觉映射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。