解决文本编辑中物体视觉一致性问题,提升生成图像保真度。
DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models

- 构建ABO-Edit数据集,含1.2万组高质量三元组图文样本。
- 发现条件嵌入空间在高噪声下仍能预测最终图像结果。
- 提出FlowMirror无参损失,不改架构即可提升生成质量。
尽管文本引导图像编辑取得显著进展,生成模型常无法保持视觉对象的一致性,即编辑过程中主体关键属性的保留。本文通过三项贡献解决此问题:首先,提出ABO-Edit数据集,专为研究对象一致性设计,包含超过12,000个三元组——源图像、编辑提示与由艺术家设计3D资产渲染的高质量目标图像,具备多视角覆盖和人工验证的质量控制。其次,揭示图像编辑修正流模型中一个被忽视的特性:条件嵌入空间虽未直接监督训练,但在高噪声水平下仍编码了对最终生成图像的预测。第三,基于该发现,提出FlowMirror,一种无需参数的辅助损失,用于监督该条件嵌入空间。无需改变模型结构,本方法在多个指标上优于基线模型,显著提升生成质量。
原文摘要 · Abstract (English)
Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions. First, we introduce ABO-Edit, a dataset specifically designed to study object consistency, comprising over 12,000 triplets of source images, editing prompts, and high-quality target images rendered from artist-designed 3D assets, with multi-view coverage and human-verified quality control. Second, we uncover an overlooked property of image-editing rectified flow models: the conditioning embedding space, not directly supervised during training, encodes a prediction of the final generated image even at high noise levels. Third, exploiting this finding, we propose FlowMirror, a parameter-free auxiliary loss that supervises this conditioning embedding space. Without architectural changes, our method improves generation quality across several metrics over baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。