arXiv:2606.08091cs.CV2026-06被引 2

让智能体自主构建视频生成流程,自动评估并优化技能。

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

论文配图:VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation
图 1 · 摘自论文原文
  • 智能体自主组合基础技能生成长视频,不依赖固定流程。
  • 技能演化使视频质量显著提升,跨框架表现差异明显。
  • 支持多模态输入,适合研究长视频生成与智能体能力评估。

近期的智能体框架(如Claude Code、Codex、OpenClaw)在工具调用和编排方面表现强劲,但其在长视频生成——这一长时程多模态任务中的能力仍缺乏深入探索。与早期基于手工设计流程的视频智能体不同,这些框架可自主构建并优化工作流。我们提出VideoWeaver,一个用于评估与演化长视频生成技能的智能体枢纽与基准测试集,其中智能体将单一指令转化为长视频,通过自定义组合基础技能构建专属工作流。该基准包含16个任务类别、285个案例,参考数据涵盖文本、图像、音频、视频及其组合。由于错误可能出现在任一阶段而非仅最终视频,我们提出“智能体作为评判者”,通过检查执行轨迹与最终视频,结合元数据和中间文件提供证据评分。基于此反馈,我们设计了技能演化算法,实现技能的精炼与合并。在多个框架与模型上实验表明:显式组合技能优于仅使用基础技能;技能演化进一步提升输出质量;性能在不同枢纽与模型间存在显著差异。所提智能体评判方法与人工判断高度一致,尤其在过程指标上表现优异。代码与数据集已开源。

原文摘要 · Abstract (English)

Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video generation, a long-horizon multimodal task, remains underexplored. Unlike earlier video agents whose pipeline is handcrafted, these frameworks can build and refine their own workflows. We introduce VideoWeaver, an agent harness and benchmark that evaluates and evolves skills for long video generation, where an agent turns a single instruction into a long video by composing foundation skills into its own workflow rather than following a predefined pipeline. The benchmark has 16 task categories and 285 cases, with references spanning text, image, audio, video, and their combinations. Because errors can arise at any stage and not just in the final video, we propose an agent-as-judge that inspects both the execution trace and the final video, grounding its scores in evidence such as metadata and intermediate files. Using this feedback, we further design a skill evolution algorithm that refines and merges the agent's skills. Across multiple frameworks and models, we find that an explicit composition skill improves the generation process over using foundation skills alone, that skill evolution further improves output quality, and that performance varies notably across harness and model choices. The proposed agent-as-judge also aligns well with human judgments, especially on process metrics. Code and dataset is available at https://github.com/JianhuiWei7/VideoWeaver

视频生成智能体技能演化多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。