arXiv:2602.01851cs.CV2026-02被引 8

构建视觉指令图像编辑基准,测试模型理解手绘等视觉提示的能力。

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

  • 设计三层次交互框架,覆盖指代定位到因果推理
  • 17个模型评测显示闭源模型领先但复杂任务表现下降
  • 专设评分框架支持细粒度评估,适合研究多模态生成的学者

近期生成模型在图像编辑方面取得显著进展,但现有系统和评测仍以文本引导为主。相比之下,人类交流本质是多模态的,如草图能高效传达空间与结构意图。为弥补这一差距,我们提出VIBE——一个包含三级交互层级的视觉指令图像编辑基准,涵盖指代定位、形态操作和因果推理。我们精心构建了反映逐步增加复杂度的高质量多样测试用例。同时,提出基于大模型的评判框架(LMM-as-a-judge)与任务特定指标,实现可扩展、细粒度的评估。对17个代表性开源与闭源图像编辑模型的全面评测表明,闭源模型具备初步视觉指令跟随能力,且持续优于开源模型;但即使最强系统在任务难度提升时性能也明显下降,揭示未来研究的重要方向。

原文摘要 · Abstract (English)

Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In contrast, human communication is inherently multimodal, where visual instructions such as sketches efficiently convey spatial and structural intent. To address this gap, we introduce VIBE, the Visual Instruction Benchmark for Image Editing with a three-level interaction hierarchy that captures deictic grounding, morphological manipulation, and causal reasoning. Across these levels, we curate high-quality and diverse test cases that reflect progressively increasing complexity in visual instruction following. We further propose a robust LMM-as-a-judge evaluation framework with task-specific metrics to enable scalable and fine-grained assessment. Through a comprehensive evaluation of 17 representative open-source and proprietary image editing models, we find that proprietary models exhibit early-stage visual instruction-following capabilities and consistently outperform open-source models. However, performance degrades markedly with increasing task difficulty even for the strongest systems, highlighting promising directions for future research.

图像编辑多模态基准评测视觉指令

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。