arXiv:2411.12293cs.CVcs.HC2024-11被引 5

用自然语言控制视频时间线编辑,让普通人也能轻松操作。

Generative Timelines for Instructed Visual Assembly

  • 构建生成式模型,理解语言指令并操控视觉时间线。
  • 自动构建数据集,无需人工标注即可训练模型。
  • 在真实场景中表现优于GPT-4o等基线模型。

本文旨在通过自然语言指令操控视觉时间线(如视频),使非专业甚至残障用户也能完成复杂的编辑任务。该任务被称为‘指令化视觉组装’,挑战在于:(i) 从输入时间线及视频库中识别并检索相关视觉内容;(ii) 理解自然语言指令;(iii) 执行所需编辑以生成目标时间线。为此,我们提出Timeline Assembler——一个用于执行指令化视觉组装的生成式模型。主要贡献有三:第一,开发了一个大型多模态语言模型,可处理视觉内容、紧凑表示时间线并准确解析编辑指令;第二,提出一种自动生成视觉组装数据集的新方法,实现无需人工标注的高效训练;第三,构建两个新数据集(图像与视频组装),验证了该模型在多种现实场景下显著优于现有基线模型(包括最新的GPT-4o),能更精准执行复杂组装指令。

原文摘要 · Abstract (English)

The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task Instructed visual assembly. This task is challenging as it requires (i) identifying relevant visual content in the input timeline as well as retrieving relevant visual content in a given input (video) collection, (ii) understanding the input natural language instruction, and (iii) performing the desired edits of the input visual timeline to produce an output timeline. To address these challenges, we propose the Timeline Assembler, a generative model trained to perform instructed visual assembly tasks. The contributions of this work are three-fold. First, we develop a large multimodal language model, which is designed to process visual content, compactly represent timelines and accurately interpret timeline editing instructions. Second, we introduce a novel method for automatically generating datasets for visual assembly tasks, enabling efficient training of our model without the need for human-labeled data. Third, we validate our approach by creating two novel datasets for image and video assembly, demonstrating that the Timeline Assembler substantially outperforms established baseline models, including the recent GPT-4o, in accurately executing complex assembly instructions across various real-world inspired scenarios.

视频生成自然语言控制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。