arXiv:2502.07096cs.HCcs.CV2025-02被引 21

用生成脚本+自动匹配片段,一键把长视频变短视频

Lotus: Creating Short Videos From Long Videos With Abstractive and Extractive Summarization

  • 先生成短文案和语音,再匹配原视频片段
  • 用户研究显示效率比手动剪辑快40%以上
  • 适合内容创作者快速生产短视频

短视频在TikTok、Instagram等平台广受欢迎,因其能迅速吸引观众注意力。许多创作者希望将长视频转为短视频,但普遍反映规划、提取和排列片段过程困难。当前做法分为两类:一是直接截取长视频片段(抽取式),保持音画同步;二是为画面添加新配音(抽象式),灵活性高但脱离原素材。本文提出Lotus系统,融合两种方式:先生成短剧本及对应语音,再将长视频片段自动匹配到生成的叙述中。创作者可借助自动化方法或编辑界面添加片段并进一步优化。与纯抽取式基线相比,用户研究发现使用Lotus制作短视频更高效,且显著优于现有工作流程。

原文摘要 · Abstract (English)

Short-form videos are popular on platforms like TikTok and Instagram as they quickly capture viewers' attention. Many creators repurpose their long-form videos to produce short-form videos, but creators report that planning, extracting, and arranging clips from long-form videos is challenging. Currently, creators make extractive short-form videos composed of existing long-form video clips or abstractive short-form videos by adding newly recorded narration to visuals. While extractive videos maintain the original connection between audio and visuals, abstractive videos offer flexibility in selecting content to be included in a shorter time. We present Lotus, a system that combines both approaches to balance preserving the original content with flexibility over the content. Lotus first creates an abstractive short-form video by generating both a short-form script and its corresponding speech, then matching long-form video clips to the generated narration. Creators can then add extractive clips with an automated method or Lotus's editing interface. Lotus's interface can be used to further refine the short-form video. We compare short-form videos generated by Lotus with those using an extractive baseline method. In our user study, we compare creating short-form videos using Lotus to participants' existing practice.

视频生成自动化剪辑内容创作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。