让视频编辑模型学会自我反思,提升推理能力与生成质量。
Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning
- 用内部视觉语言模型做自我评估,实现可微分反馈优化编辑过程。
- 在RVE-Bench上,编辑准确率和画质提升10%,优于微调基线。
- 适合研究视频生成、推理增强与自洽学习的开发者与学者。
统一视频模型虽具备强大的理解与生成能力,但在需推理支持的视觉编辑任务中仍表现不佳,尤其当其内置视觉-语言模型(VLM)强大时。我们归因于两点:(1)现有数据集不足以训练和评估推理感知的视频编辑;(2)模型的推理与编辑能力之间存在固有断层,导致理解无法有效指导编辑。为此,我们提出“推理感知视频编辑”(RVE)任务,要求在编辑中推理物理合理性与因果动态。为系统评估,构建了涵盖两个互补子集的基准RVE-Bench:推理感知视频编辑(RAVE)与上下文视频到视频生成(ICVG),覆盖多种推理维度。在此基础上,提出ReViSE框架,利用模型内建VLM作为自我反思评估器,在训练中对生成结果进行评估并反向优化,无需外部评判者。实验表明,该方法在RAVE子集上相较微调模型整体得分提升10%,验证了可微分自我反馈的有效性。
原文摘要 · Abstract (English)
Unified video models exhibit strong capabilities in understanding and generation, yet they struggle with reason-informed visual editing even when equipped with powerful internal vision-language models (VLMs). We attribute this gap to two factors: (1) existing datasets are inadequate for training and evaluating reasoning-aware video editing, and (2) an inherent disconnect between the models' reasoning and editing capabilities, which prevents understanding from guiding the editing process. To address this, we introduce the Reason-Informed Video Editing (RVE) task, which requires reasoning about physical plausibility and causal dynamics during editing. To support systematic evaluation, we construct RVE-Bench, a comprehensive benchmark with two complementary subsets: Reasoning-Aware Video Editing (RAVE) and In-Context Video-to-Video Generation (ICVG), spanning diverse reasoning dimensions across both editing and generation scenarios. Building upon this foundation, we propose ReViSE, a self-reflective learning framework that harnesses the model's internal VLM to evaluate and refine its own generation during training. Unlike prior reward-based approaches that rely on external critics, ReViSE leverages the model's internal VLM as a self-reflective evaluator, providing differentiable feedback that directly refines the generator's reasoning behavior during training. Extensive experiments on RVE-Bench demonstrate that ReViSE enhances editing accuracy and visual fidelity, outperforming the finetuned counterpart by 10% in Overall score on the RAVE subset, demonstrating the effectiveness of self-reflective differentiable reward.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。