arXiv:2605.30437cs.CV2026-05

让AI修图既好看又不失真,修复图像结构错位与虚构内容。

Mitigating Content Shift and Hallucination in GenAI Image Editing via Structural Refinement

论文配图:Mitigating Content Shift and Hallucination in GenAI Image Editing via Structural Refinement
图 1 · 摘自论文原文
  • 先对齐原图与AI生成图的结构和光影,再融合增强效果。
  • 在保持原始分辨率的前提下,显著减少纹理扭曲和幻觉内容。
  • 适合需要精准像素级输出的图像编辑场景,如医疗或设计领域。

生成式AI图像编辑器(如Nano Banana)能通过文本提示实现直观的图像修饰,使非专业人士也能轻松操作。然而,生成模型常导致空间错位、纹理失真和内容幻觉,影响后续需要像素级一致性的下游任务。本文提出一种“结构保持型生成式图像融合”问题,目标是在保留生成结果视觉美感的同时,确保其结构与原始输入一致。为此,我们设计了一种后处理框架:首先建立输入图与生成图之间的粗略空间与光照对应关系,再通过融合阶段转移期望的增强效果,同时抑制幻觉内容。由于该设定缺乏直接先例,我们在无监督图像融合与照片级风格迁移方法上进行对比实验。结果表明,本方法在保持美学质量的同时,更有效地维持了像素级结构一致性与输入分辨率。

原文摘要 · Abstract (English)

Generative AI (GenAI) image editors, such as Nano Banana, produce visually compelling results for retouching tasks, enabling non-experts to edit images through text prompts alone. However, the generative nature of these models often introduces spatial misalignment, texture distortion, and content hallucination, all of which are detrimental to downstream workflows that require pixel-level fidelity. We identify a problem setting we call "structure-preserving GenAI fusion" for black-box GenAI image retouching: retain the perceptual enhancements of a GenAI output while enforcing structural faithfulness to the original input image. To address this problem, we propose a post-processing framework that fuses an input image with its GenAI-enhanced counterpart by first establishing coarse spatial and photometric correspondences, then performing a fusion stage that transfers desired enhancements while suppressing hallucinated content. In the absence of direct prior work in this setting, we evaluate our framework against representative methods from photorealistic style transfer and image fusion. Our experiments demonstrate that our method better preserves aesthetic quality while maintaining pixel-level structural consistency and the input resolution.

图像编辑生成模型结构保持幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。