arXiv:2608.16328cs.CV2026-08被引 1

用二进制证据高效建模视频编辑意图,轻量级框架性能超越大模型。

GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks

论文配图:GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks
图 1 · 摘自论文原文
  • 将编辑意图转化为比特级保留或翻转决策,以二进制证据驱动生成。
  • 仅用0.6M数据和3%参数量,2B模型性能超多个14B开源编辑器。
  • 适合追求高效、低资源视频编辑的开发者与研究者使用。

基于指令的通用视频编辑旨在统一多样化的编辑操作于单一直观界面。现有方法常依赖高资源条件机制,如重型分支或昂贵的源图像拼接。本文提出GRNEdit,一种轻量级两阶段框架。GRN通过比特组合编码视觉语义,经任务特定微调后,将编辑语义重新定义为对单个比特的局部保留或翻转决策。源信息被建模为支持观测二进制状态的坐标式证据,而GRN主干负责将这些局部决策整合为连贯生成语义。第一阶段中,紧凑编码器将离散源码转换为连续证据信号,由GRN在二进制精炼过程中吸收。受无提示训练启发,我们为零条件赋予特定编辑意义:空指令表示无编辑,通过源重建进行监督。该身份路径不仅隐式增强证据利用与内容保留,还生成与编辑态同空间的源保持状态。第二阶段可直接比较编辑态与其源保持态的差异,修正未决的目标比特决策。在仅0.6M样本、条件参数占比不足3%的情况下,GRNEdit-2B与GRNEdit-8B在OpenVE-Bench上分别取得4.03与4.18分,2B模型优于多个14B开源编辑器,8B模型性能媲美领先开源方案。

原文摘要 · Abstract (English)

Instruction-based general video editing seeks to unify diverse editing operations within a single, intuitive interface. Existing approaches often rely on resource-intensive conditioning, using either heavyweight branches or costly source concatenation. Is there any efficient way to model editing intent? Thus, we introduce GRNEdit, a lightweight two-stage framework. GRN inspires our approach by encoding visual semantics through combinations of bits. Through task-specific fine-tuning, we take this representation further and recast editing semantics as local retain-or-flip decisions over individual bits. Source information is consequently modeled as coordinate-wise evidence supporting the observed binary states, while the GRN backbone remains responsible for resolving their global composition into coherent generative semantics. In Stage I, a compact encoder translates discrete source codes into continuous evidence signals, which GRN assimilates throughout binary refinement. Inspired by null-prompt training for classifier-free guidance, we further assign the null condition an editing-specific meaning: an empty instruction denotes no edit and is supervised through source reconstruction. This identity pathway not only implicitly strengthens evidence utilization and content preservation in Stage I, but also produces a source-preserving state in the same representation space as the edited state. Stage II can therefore directly compare each edited state with its source-preserving counterpart and use their discrepancy to revise unresolved target-bit decisions. Trained on only 0.6M pairs with less than 3\% conditioning parameters, GRNEdit-2B and GRNEdit-8B achieve scores of 4.03 and 4.18 on OpenVE-Bench. The 2B model outperforms multiple 14B open-source editors, while the 8B model performs on par with leading open-source editors.

视频编辑生成模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。