构建可扩展的精准图像编辑评测基准,解决模型在精确操作上的短板。
PaintBench: Deterministic Evaluation of Precise Visual Editing

- 设计动态可扩展的基准,覆盖20种精准图像编辑任务。
- 最高模型仅达17.1% mIoU,凸显精准编辑仍存巨大挑战。
- 适合研究视觉编辑鲁棒性与泛化能力的学者使用。
当前多模态模型在开放式图像编辑上表现良好,但在精确的单答案编辑任务中仍面临重大挑战。为此,我们提出PaintBench,一个动态可扩展的基准,涵盖四类共20种基础精准图像编辑操作:几何变换、结构操作、色彩修改和符号推理。通过可配置复杂度的程序生成,实现无限且无污染的评估集,采用确定性的像素级评估,避免依赖有偏的评分模型。对11个图像编辑模型的测试显示整体性能较低,当前最高性能工业模型仅得17.1%(mIoU)。任务分解揭示几何变换、多数结构操作及基于公式的色彩修改尤为困难,且模型存在特定专长。细粒度诊断表明,物体数量、背景复杂度、色彩方案和编辑区域大小均导致性能下降。为验证得分的泛化性,我们构建了程序化、确定性的数据可视化编辑评估(TinyGrafixBench),发现其与PaintBench得分呈强线性相关(R² = 0.91,p < 0.001)。总体而言,PaintBench为精准多模态视觉编辑的测量与进展提供了严谨基础。
原文摘要 · Abstract (English)
While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle. To probe this challenge, we introduce PaintBench, a dynamically scalable benchmark targeting 20 fundamental precise visual editing operations across four categories: geometric transformation, structural manipulation, color change, and symbolic reasoning. Procedural generation with configurable complexity enables an effectively infinite, contamination-resistant evaluation suite, and deterministic pixel-level evaluation eliminates reliance on bias-prone judge models. Across 11 image editing models, we find overall low performance, with the current highest-performing industry leader scoring only 17.1% (mIoU). Task decomposition reveals especially challenging operation types (geometric transformation, most structural manipulation, formula-based color change) and model-specific specializations. Fine-grained benchmark diagnostics further show performance degradations induced by scene variations in object count, background complexity, color scheme, and edit-region size. To test generalization of PaintBench scores to applied task performance, we create a procedural, deterministic evaluation for data visualization editing (TinyGrafixBench) and find strong linear correlation with PaintBench scores ($R^2 = 0.91$, $p < 0.001$). Altogether, PaintBench provides a rigorous foundation for measuring and driving progress in precise multimodal visual editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。