让视频生成模型像导演一样自我改进,持续优化长故事创作流程。
CineForge: Self-Improving Agents for Long-Horizon Video Generation

- 用多状态分解故事,自动协调拍摄与生成过程
- 跨故事学习失败模式,使生成质量从4.024提升至4.380
- 适合需要长期叙事生成的智能内容生产场景
长时序叙事视频生成需生产代理协调情节拆解、状态追踪、镜头设计、提示构建、渲染与修订等跨场景任务。现有自适应系统多聚焦请求或技能微调,未能将重复性生产失败转化为持续性的阶段性改进。本文提出CineForge,一个自演化视频生产框架,包含CineForge-Produce(生成)与CineForge-Evolve(进化)两部分。CineForge-Produce将源故事分解为类型化的情节、角色、空间与影视状态,据此协调资产与片段生成,并记录为标准生产轨迹。CineForge-Evolve采用案例-模式-策略演化(CPPE)机制,分析轨迹证据,将共现问题归纳为局部阶段修复补丁,并通过结构重播与置信度控制的成对评估部署更新。为衡量完整故事实现度,引入CineScope:包含100个脚本的CineScope-Data数据集,以及涵盖因果状态、导演调度、节奏分配与角色弧线的人类对齐多尺度指标。在CineScope-Data及两个公开基准上,演进后的CineForge策略使CineScope-Metric从4.024提升至4.380,优于三项长视频基线,且在新故事上减少37.0%的审查大模型调用次数。结果证明,生产轨迹可作为视频代理的可行动经验,实现长叙事任务中的累积性改进。
原文摘要 · Abstract (English)
Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt construction, rendering, and revision across interdependent scenes. Existing adaptive video systems primarily refine requests or reusable skills, leaving recurring production failures disconnected from persistent, stage-targeted improvements across stories. We introduce CineForge, a self-evolving video-production agent framework that couples CineForge-Produce for video generation with CineForge-Evolve for cross-story policy evolution. CineForge-Produce organizes each source story into typed narrative, character, spatial, and cinematic states, uses them to coordinate asset and clip generation, and records the process as a canonical production trajectory. CineForge-Evolve applies Case-to-Pattern-to-Policy Evolution (CPPE) to review trajectory evidence, consolidate recurrent findings into bounded stage-local patches, and deploy validated updates through structural replay and confidence-controlled paired evaluation. To measure complete story realization, we introduce CineScope, which combines a 100-script CineScope-Data suite with a human-aligned, multiscale CineScope-Metric spanning causal state, directorial orchestration, pacing and resource allocation, and character arc. Across CineScope-Data and two public benchmarks, the evolved CineForge policy improves CineScope-Metric from 4.024 to 4.380, outperforms three long-video baselines with consistent gains under ScriptAgent, and reduces review LLM calls by 37.0% on new stories. These results establish production trajectories as actionable experience for video agents that improve cumulatively across long-form storytelling tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。