arXiv:2605.14842cs.CV2026-05

首个专注抽象图像编辑的基准,让模型理解'氛围'等模糊指令

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

论文配图:Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
图 1 · 摘自论文原文
  • 将抽象编辑拆解为实体级评估,用新框架衡量模型表现
  • 11个主流模型在抽象指令上普遍过度或不足编辑,效果不佳
  • 结合大模型编码器与迭代思考可显著提升抽象意图理解能力

人类自然使用'情绪'等抽象概念交流,但现有图像编辑评测多聚焦显式指令,忽视抽象要求。本文首次形式化抽象图像编辑的定义与分类体系,提出Entity-Rubrics框架,将抽象编辑分解为个体实体层面的评估,与人类判断高度相关。同时构建AbstractEdit——首个面向多样化真实场景的抽象图像编辑基准。在该数据集上评估11个领先模型发现:标准架构难以平衡意图执行与内容保留,常出现编辑不足或过度。分析表明,提升性能关键在于融合先进LLM文本编码器与迭代推理机制。未来,该实体范式可扩展为奖励模型、帮助模型理解抽象表达,或用于测试时错误诊断。本工作旨在推动多模态交互迈向更自然、开放的人机沟通。

原文摘要 · Abstract (English)

Humans naturally communicate through abstract concepts like "mood". However, current image editing benchmarks focus primarily on explicit, literal commands, leaving abstract instructions largely underexplored. In this work, we first formalize the definition and taxonomy of abstract image editing. To measure instruction-following in this challenging domain, we introduce Entity-Rubrics, a framework that breaks down abstract edits into individual, entity-level assessments and achieves strong correlation with human judgment. Alongside this framework, we contribute AbstractEdit, the first benchmark dedicated to abstract image editing across diverse real-world scenes. Evaluating 11 leading models on this dataset reveals a fundamental challenge: standard architectures struggle to balance intent and preservation, commonly defaulting to under-editing or over-editing. Our analysis demonstrates that driving meaningful improvements relies heavily on integrating advanced LLM text encoders and iterative thinking. Looking forward, our entity-based paradigm can generalize beyond assessment to serve as a reward model, enable models to correctly interpret abstract communication, or highlight specific failures in test-time critique loops. Ultimately, we hope this work serves as a stepping stone toward seamless multimodal interaction, closing the gap between rigid machine execution and the natural, open-ended way humans communicate.

图像编辑抽象意图多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。