让AI更懂用户创作意图,生成更专业的可控视频
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition

- 用动画制作数据训练专用视觉语言模型,精准理解模糊创意
- 通过强化学习对齐推理输出,实现多条件统一控制
- 支持专业场景如分镜草图生成,适合影视动画创作者
当前扩散模型在视频生成中具备强逼真度和流畅性,但在抽象、稀疏或复杂条件下仍表现脆弱,难以满足分镜草图、黏土渲染等专业生产流程需求。现有模型依赖适配器或通用视觉语言模型(VLM),难以准确捕捉用户创作意图。我们提出CogOmniControl,一个基于推理驱动的可控视频生成框架,将任务分解为创作意图认知与生成两部分。具体地,使用真实动漫制作数据训练专用CogVLM,相比通用VLM,能更专业清晰地理解稀疏抽象输入,并生成密集推理输出。同时,CogOmniDiT通过上下文生成统一多种控制信号,并经强化学习与CogVLM输出对齐。利用CogVLM的引导能力,我们设计特定评估器并实现Best-of-N选择,形成闭环‘操纵杆式’架构。我们构建了基于真实创作流程的CogReasonBench和CogControlBench,实验表明其优于现有开源模型。
原文摘要 · Abstract (English)
Recent diffusion models achieve strong photorealism and fluency in video generation, yet remain fragile under abstract, sparse or complex conditions, leading to poor performance in professional production workflows such as storyboard sketches and clay render conditions. Existing video generation models, either inject conditions through adapters or couple a generic vision-language model (VLM) within a diffusion backbone, leaving a capability gap and failing to produce the videos that align with the user's creative intent. We present CogOmniControl, a reasoning-driven framework that factorizes controllable video generation into creative intent cognition and generation. Specifically, we train a specialized CogVLM using authentic anime production data. Compared to generic VLMs, it generates more professional and clear outputs, accurately cognizing user creative intent from sparse and abstract conditions and tuning these cues into dense reasoning output. Besides, CogOmniDiT unifies the controls from various conditions through in-context generation and is aligned to the CogVLM reasoning outputs via reinforcement learning. Furthermore, leveraging CogVLM's robust capability in guiding video generation, we release its potential in planning specific evaluators and enable a Best-of-N selection for the generated videos. This integration transforms the entire framework into a closed-loop "harness-like" architecture. We further introduce CogReasonBench and CogControlBench, built from professional workflows data that carry genuine creative intent rather than simulated ones. Experiments on two benchmarks show that CogOmniControl surpassed the existing open-source models. The project website: https://um-lab.github.io/CogOmniControl/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。