arXiv:2505.12237cs.CV2025-05被引 8

用语言描述统一视频剪辑,让大模型更懂故事逻辑。

From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations

  • 将视频片段转为语言结构化描述,适配大模型处理
  • 提出新策略提升剪辑顺序的逻辑一致性和稳定性
  • 适合智能视频编辑、内容创作方向的研究者

大型语言模型(LLMs)和视觉-语言模型(VLMs)在视频理解中展现出强大推理与泛化能力,但其在视频编辑中的应用仍不充分。本文首次系统研究了LLMs在视频编辑中的作用。为弥合视觉信息与语言推理之间的鸿沟,我们提出L-Storyboard,一种将离散视频片段转化为适合大模型处理的结构化语言描述的中间表示。我们将视频编辑任务分为收敛型与发散型,聚焦三项核心任务:镜头属性分类、下一镜头选择与镜头序列排序。针对发散任务输出不稳定的问题,提出StoryFlow策略,将多路径推理过程转化为收敛式选择机制,显著提升任务准确率与逻辑连贯性。实验表明,L-Storyboard增强了视觉信息与语言描述间的稳健映射,大幅提高任务可解释性与隐私保护能力;StoryFlow则有效提升了镜头序列排序的逻辑一致性与输出稳定性,凸显了大模型在智能视频编辑中的巨大潜力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable reasoning and generalization capabilities in video understanding; however, their application in video editing remains largely underexplored. This paper presents the first systematic study of LLMs in the context of video editing. To bridge the gap between visual information and language-based reasoning, we introduce L-Storyboard, an intermediate representation that transforms discrete video shots into structured language descriptions suitable for LLM processing. We categorize video editing tasks into Convergent Tasks and Divergent Tasks, focusing on three core tasks: Shot Attributes Classification, Next Shot Selection, and Shot Sequence Ordering. To address the inherent instability of divergent task outputs, we propose the StoryFlow strategy, which converts the divergent multi-path reasoning process into a convergent selection mechanism, effectively enhancing task accuracy and logical coherence. Experimental results demonstrate that L-Storyboard facilitates a more robust mapping between visual information and language descriptions, significantly improving the interpretability and privacy protection of video editing tasks. Furthermore, StoryFlow enhances the logical consistency and output stability in Shot Sequence Ordering, underscoring the substantial potential of LLMs in intelligent video editing.

视频编辑大模型语言表示逻辑排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。