arXiv:2501.06173cs.CV2025-01ICCV被引 19

构建烹饪长视频数据集,提升生成视频的叙事连贯性

VideoAuteur: Towards Long Narrative Video Generation

  • 设计大规模烹饪视频数据集,支持长序列叙事生成
  • 新方法使关键帧视觉细节与语义更匹配,提升整体质量
  • 适合研究长视频生成、跨模态对齐的学者使用

近期视频生成模型在生成数秒高质量视频片段方面取得进展,但在生成能清晰传达事件的长序列视频方面仍面临挑战,限制了其在连贯叙述中的应用。本文提出一个大规模烹饪视频数据集,用于推动烹饪领域长篇叙事生成的发展。通过先进视觉语言模型(VLMs)评估图像保真度,利用视频生成模型验证文本描述准确性,验证了数据集质量。我们进一步引入长叙事视频导演(Long Narrative Video Director),通过对齐视觉嵌入,增强生成视频的视觉与语义一致性。该方法结合文本与图像嵌入的微调技术,在生成更具视觉细节和语义一致性的关键帧方面表现显著提升。项目页面:https://videoauteur.github.io/

原文摘要 · Abstract (English)

Recent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long sequences that convey clear and informative events, limiting their ability to support coherent narrations. In this paper, we present a large-scale cooking video dataset designed to advance long-form narrative generation in the cooking domain. We validate the quality of our proposed dataset in terms of visual fidelity and textual caption accuracy using state-of-the-art Vision-Language Models (VLMs) and video generation models, respectively. We further introduce a Long Narrative Video Director to enhance both visual and semantic coherence in generated videos and emphasize the role of aligning visual embeddings to achieve improved overall video quality. Our method demonstrates substantial improvements in generating visually detailed and semantically aligned keyframes, supported by finetuning techniques that integrate text and image embeddings within the video generation process. Project page: https://videoauteur.github.io/

视频生成长视频叙事生成多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。