arXiv:2512.12196cs.MMcs.CV2025-12被引 3

AutoMV用多智能体自动生成结构完整、连贯的音乐视频。

AutoMV: An Automatic Multi-Agent System for Music Video Generation

  • 通过多智能体协作,解析音乐结构并生成分镜脚本。
  • 在四个评估维度上超越现有基线,接近专业人工制作水平。
  • 适合需要快速生成高质量音乐视频的创作者或团队使用。

全长歌曲的音乐-视频(M2V)生成面临诸多挑战:现有方法仅能生成短而零散的片段,难以与音乐结构、节拍或歌词对齐,且缺乏时间一致性。本文提出AutoMV,一种直接从歌曲生成完整音乐视频(MV)的多智能体系统。该系统首先利用音乐处理工具提取结构、人声轨道和时间对齐的歌词等属性,并作为后续智能体的上下文输入。编剧智能体与导演智能体基于这些信息设计短剧本,定义角色档案并存入共享外部数据库,同时指定镜头指令。随后,各智能体调用图像生成器生成关键帧,以及不同视频生成器完成“故事”或“歌手”场景。验证智能体评估输出结果,实现多智能体协同,生成连贯的长视频。为进一步评估M2V生成效果,我们构建了一个包含四大类(音乐内容、技术、后期、艺术)和十二项细粒度指标的基准测试。该基准用于对比商用产品、AutoMV与人工执导的MV,由专家人类评分者打分:AutoMV在所有四类中均显著优于现有基线,缩小了与专业作品的差距。最后,我们探索使用大型多模态模型作为自动评分工具;虽具潜力,但仍逊于人类专家,表明未来仍有提升空间。

原文摘要 · Abstract (English)

Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We propose AutoMV, a multi-agent system that generates full music videos (MVs) directly from a song. AutoMV first applies music processing tools to extract musical attributes, such as structure, vocal tracks, and time-aligned lyrics, and constructs these features as contextual inputs for following agents. The screenwriter Agent and director Agent then use this information to design short script, define character profiles in a shared external bank, and specify camera instructions. Subsequently, these agents call the image generator for keyframes and different video generators for "story" or "singer" scenes. A Verifier Agent evaluates their output, enabling multi-agent collaboration to produce a coherent longform MV. To evaluate M2V generation, we further propose a benchmark with four high-level categories (Music Content, Technical, Post-production, Art) and twelve ine-grained criteria. This benchmark was applied to compare commercial products, AutoMV, and human-directed MVs with expert human raters: AutoMV outperforms current baselines significantly across all four categories, narrowing the gap to professional MVs. Finally, we investigate using large multimodal models as automatic MV judges; while promising, they still lag behind human expert, highlighting room for future work.

音乐视频多智能体自动生成内容生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。