arXiv:2511.20563cs.CV2025-11

让AI更懂用户指令,精准生成可控视频

A Reason-then-Describe Instruction Interpreter for Controllable Video Generation

  • 先分析指令意图,再生成详细可控的生成指引
  • 在多场景下提升指令匹配度与视频生成质量
  • 适合需要精准控制视频内容的研究与应用

扩散变换模型显著提升了视频保真度和时序连贯性,但实际可控性仍受限。用户输入常简洁、模糊且组合复杂,与训练中使用的详细提示不匹配,导致意图与输出不符。我们提出 ReaDe,一种通用、模型无关的指令解释器,可将原始指令转化为下游视频生成器所需的精确操作规范。ReaDe 采用‘先推理后描述’范式:首先分析用户请求,识别核心需求并化解歧义,然后生成详尽指导,实现忠实可控的生成。通过两阶段优化训练:(i) 增强推理的监督提供分步推理轨迹和密集标注;(ii) 多维奖励分配器实现稳定、反馈驱动的自然风格描述优化。在单条件与多条件场景下的实验表明,ReaDe 在指令忠实度、描述准确性和下游视频质量上均有持续提升,并展现出对高推理强度及未见输入的强大泛化能力。ReaDe 为对齐可控视频生成与用户真实意图提供了可行路径。

原文摘要 · Abstract (English)

Diffusion Transformers have significantly improved video fidelity and temporal coherence, however, practical controllability remains limited. Concise, ambiguous, and compositionally complex user inputs contrast with the detailed prompts used in training, yielding an intent-output mismatch. We propose ReaDe, a universal, model-agnostic interpreter that converts raw instructions into precise, actionable specifications for downstream video generators. ReaDe follows a reason-then-describe paradigm: it first analyzes the user request to identify core requirements and resolve ambiguities, then produces detailed guidance that enables faithful, controllable generation. We train ReaDe via a two-stage optimization: (i) reasoning-augmented supervision imparts analytic parsing with stepwise traces and dense captions, and (ii) a multi-dimensional reward assigner enables stable, feedback-driven refinement for natural-style captions. Experiments across single- and multi-condition scenarios show consistent gains in instruction fidelity, caption accuracy, and downstream video quality, with strong generalization to reasoning-intensive and unseen inputs. ReaDe offers a practical route to aligning controllable video generation with accurately interpreted user intent. Project Page: https://sqwu.top/ReaDe/.

视频生成指令理解可控生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。