新基准TECCI挑战文本编辑模型在复杂指令下的表现极限。
TECCI: Tricky Edits of Collected and Curated Images

- 构建7类图像与530张人工设计难题,覆盖位置、运动等5类编辑类型
- 五模型平均成功率不足22%,纳米香蕉专业版最优
- 空间布局与创意编辑最难,色彩修改最易,可自动评测
尽管近期进展显著,当前文本引导图像编辑方法在遵循指令、最小化修改和保证视觉质量方面仍面临挑战,尤其在涉及位置、运动、视角、尺度及创意编辑等复杂操作时。为系统评估生成式图像编辑器,我们提出新基准TECCI:收集与筛选的图像难题集。TECCI包含7个图像类别,共7550对图像与编辑指令。指令由Gemini自动生成,涵盖每张图5种编辑类型,并额外收录530张含人工设计难题的图像。我们通过人工评估五个主流模型,从指令遵循度、修改最小性、视觉质量三维度打分;并基于Gemini构建自动评分器,准确率达74.7%。评估发现:1)所有模型整体成功率均未超过22%;2)Nano Banana Pro表现最佳;3)模型在指令遵循上优于修改最小性和视觉质量;4)建筑与自然图像编辑尤为困难,需强空间理解力;5)推理与创意编辑最难,颜色与外观修改最易。
原文摘要 · Abstract (English)
Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction following, minimally editing the source image, and ensuring high visual quality. These problems are especially apparent when the requested edit is challenging, such as those that involve position, motion, viewpoint, scale and creative edits. To systematically test generative image editors, we propose a novel image editing benchmark -- TECCI: Tricky Edits of Collected and Curated Images. TECCI consists of a completely new set of images we are releasing. The images in TECCI span 7 image categories. The images and these categories were curated intentionally to target weaknesses of existing methods. The edit instructions in TECCI are automatically generated by Gemini, covering 5 edit types per source image. We also curated a set of 530 images for which we created challenging manually written edit instructions. Overall, TECCI contains 7550 pairs of images and edit instructions. We conduct human evaluations of five leading image editing models on TECCI. Humans judge outputs along three dimensions: 1) instruction following, 2) minimality of the edits, and 3) visual quality. To scale-up the evaluation, we also build an auto-rater using Gemini that achieves 74.7% accuracy in matching human evaluations. Our evaluations reveal that: 1) none of the models exceed a 22% overall success rate, demonstrating the challenging nature of TECCI, 2) Nano Banana Pro is the best performing model overall, 3) models perform significantly better at instruction following compared to minimal edits and visual quality, 4) models struggle with editing architecture and nature images which require strong understanding of spatial layout and intricate visual details. 5) reasoning and creative edits are the most difficult, whereas color and appearance edits are the easiest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。