arXiv:2509.16811cs.AIcs.HC2025-09被引 7

用自然语言指令自动重编长篇叙事视频,保留故事连贯性。

Prompt-Driven Agentic Video Editing System: Autonomous Comprehension of Long-Form, Story-Driven Media

  • 通过提示词驱动的模块化系统,无需时间线即可重构多小时视频。
  • 在400多个视频上验证,专家评分和偏好测试均优于传统方法。
  • 支持创作者透明干预中间结果,兼顾自动化与创作控制权。

创作者编辑长篇叙事视频的难点不在于界面复杂,而在于需处理海量素材的搜索、分镜与排序认知负荷。现有基于转录或嵌入的方法难以满足创意工作流需求,因模型难以追踪角色、推断动机、连接分散事件。本文提出一种提示驱动的模块化编辑系统,让创作者通过自由文本指令重构多小时内容,而非依赖时间线。核心为语义索引管道,结合时序分割、引导式记忆压缩与跨粒度融合,构建全局叙事,生成可解释的剧情、对话、情绪与上下文痕迹。用户可获得电影级剪辑结果,并可选择性优化透明中间输出。在400多个视频上通过专家评分、问答测试与偏好研究评估,系统有效提升提示驱动编辑能力,保持叙事连贯性,平衡自动化与创作者控制。

原文摘要 · Abstract (English)

Creators struggle to edit long-form, narrative-rich videos not because of UI complexity, but due to the cognitive demands of searching, storyboarding, and sequencing hours of footage. Existing transcript- or embedding-based methods fall short for creative workflows, as models struggle to track characters, infer motivations, and connect dispersed events. We present a prompt-driven, modular editing system that helps creators restructure multi-hour content through free-form prompts rather than timelines. At its core is a semantic indexing pipeline that builds a global narrative via temporal segmentation, guided memory compression, and cross-granularity fusion, producing interpretable traces of plot, dialogue, emotion, and context. Users receive cinematic edits while optionally refining transparent intermediate outputs. Evaluated on 400+ videos with expert ratings, QA, and preference studies, our system scales prompt-driven editing, preserves narrative coherence, and balances automation with creator control.

视频编辑提示工程叙事理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。