让对话图像编辑中被遮挡的内容自动恢复,避免错误生成。
Making Implicit Preservation Intent Explicit in Conversational Image Editing

- 用历史图像和指令显式提示应保留的内容
- 在基准测试中恢复准确率提升23.6%,一致性更强
- 适合做对话式图像编辑或内容一致性研究的团队
对话式图像编辑不仅需保留可见内容,还需恢复跨轮次临时消失的区域。当新增或修改内容遮挡了原可见区域时,若该区域语义未变,应恢复显示。现有系统常无法还原此类被遮挡但未改变的内容,导致不一致或幻觉结果。本文提出OCCUR-Bench,一个用于评估时间连续性的诊断性基准,涵盖多样遮挡与重现场景,并提供历史还原参考。同时提出ReSpec,一种无需训练的框架,通过将恢复意识指令与历史视觉参考配对,使隐含的保留意图显式化。给定编辑历史,ReSpec识别应持久化的区域,选取能提供缺失视觉证据的历史图像状态,并以指令与参考图条件化上下文编辑器。实验表明,ReSpec在OCCUR-Bench上显著提升恢复保真度与时间一致性,强调应基于编辑历史而非仅当前图像来定义保留行为。
原文摘要 · Abstract (English)
Conversational image editing requires preserving not only visible content, but also content that temporarily disappears across turns. When newly added or modified content occludes a previously visible region, that region should reappear if it was never semantically changed. However, existing systems often fail to recover such occluded-but-unchanged content, producing inconsistent or hallucinated results. We introduce OCCUR-Bench, a diagnostic benchmark for temporal preservation in conversational image editing. OCCUR-Bench provides diverse occlusion-and-revelation scenarios with historical restoration references, enabling evaluation of faithful restoration rather than plausible regeneration. We also propose ReSpec, a training-free framework that makes implicit preservation explicit by pairing restoration-aware instructions with historical visual references. Given an editing history, ReSpec identifies what should persist, selects the historical image state that provides missing visual evidence, and conditions an in-context editor on the resulting instruction and reference image. Experiments show that ReSpec improves restoration fidelity and temporal consistency on OCCUR-Bench, highlighting the need to ground preservation in editing history rather than only the current image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。