AutoCut一键生成广告视频,让多模态内容高效协同。
AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

- 用统一标记空间融合视频、音频与文本特征
- 在真实数据集上降低制作成本并提升一致性
- 适合需要快速生成广告视频的创作者
短视频已成为数字广告的主要媒介,但当前内容创作流程分散且模态割裂,导致制作成本高、效率低。为此,我们提出AutoCut,一个基于多模态离散化与可控生成的端到端广告视频编辑框架。该框架通过专用编码器提取视频与音频特征,采用残差向量量化将其离散化为与文本对齐的统一标记,构建共享的视频-音频-文本标记空间。在此基础上,基于基础模型开发多模态大语言模型,通过多模态对齐与监督微调,支持视频选段排序、脚本生成与背景音乐选择等任务。最终,完整生产流水线将预测的标记序列转化为可部署的长视频输出。在真实广告数据集上的实验表明,AutoCut显著降低制作成本与迭代时间,大幅提升内容一致性和可控性,为规模化视频生成提供新路径。
原文摘要 · Abstract (English)
Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint and modality-specific, leading to high production costs and low overall efficiency. To address this issue, we propose AutoCut, an end-to-end advertisement video editing framework based on multimodal discretization and controllable editing. AutoCut employs dedicated encoders to extract video and audio features, then applies residual vector quantization to discretize them into unified tokens aligned with textual representations, constructing a shared video-audio-text token space. Built upon a foundation model, we further develop a multimodal large language model for video editing through combined multimodal alignment and supervised fine-tuning, supporting tasks covering video selection and ordering, script generation, and background music selection within a unified editing framework. Finally, a complete production pipeline converts the predicted token sequences into deployable long video outputs. Experiments on real-world advertisement datasets show that AutoCut reduces production cost and iteration time while substantially improving consistency and controllability, paving the way for scalable video creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。