构建统一视频图像编辑评估基准,用轻量模型高效替代大模型打分。
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs

- 统一图像与视频编辑评测协议,覆盖9类图像和8类视频操作。
- 轻量4B/8B模型与人类评分高度一致,成本降低超90%。
- 适合研究视觉编辑模型、追求高效评估的开发者使用。
视觉编辑模型的评估仍分散于不同方法与模态之间。现有基准多针对特定范式设计,难以公平跨范式比较,且视频编辑缺乏可靠评估标准。同时,常用自动指标常与人类偏好不符,而直接使用大型多模态模型(MLLM)作为评估器则带来高昂计算与财务成本。我们提出UniEditBench,一个支持重建型与指令驱动型方法的统一图像与视频编辑基准。该基准包含九类图像操作(添加、移除、替换、改变、笔画式、提取、调整、计数、重排)和八类视频操作,涵盖计数与空间重排等复杂组合任务。为实现可扩展评估,我们将高容量的MLLM裁判(Qwen3-VL-235B-A22B Instruct)蒸馏为轻量级4B/8B评估器,提供结构保真度、文本对齐、背景一致性、自然性及时空一致性(视频)等多维评分。实验表明,蒸馏后的评估器与人类判断高度一致,部署成本相比教师模型显著降低。UniEditBench为现代视觉编辑方法提供了实用且可复现的评测协议。相关基准与奖励模型已公开:https://github.com/wesar1/UniEditBench。
原文摘要 · Abstract (English)
The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms, making fair cross-paradigm comparisons difficult, while video editing lacks reliable evaluation benchmarks. Furthermore, common automatic metrics often misalign with human preference, yet directly deploying large multimodal models (MLLMs) as evaluators incurs prohibitive computational and financial costs. We present UniEditBench, a unified benchmark for image and video editing that supports reconstruction-based and instruction-driven methods under a shared protocol. UniEditBench includes a structured taxonomy of nine image operations (Add, Remove, Replace, Change, Stroke-based, Extract, Adjust, Count, Reorder) and eight video operations, with coverage of challenging compositional tasks such as counting and spatial reordering. To enable scalable evaluation, we distill a high-capacity MLLM judge (Qwen3-VL-235B-A22B Instruct) into lightweight 4B/8B evaluators that provide multi-dimensional scoring over structural fidelity, text alignment, background consistency, naturalness, and temporal-spatial consistency (for videos). Experiments show that the distilled evaluators maintain strong agreement with human judgments and substantially reduce deployment cost relative to the teacher model. UniEditBench provides a practical and reproducible protocol for benchmarking modern visual editing methods. Our benchmark and the associated reward models are publicly available at https://github.com/wesar1/UniEditBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。