arXiv:2503.07047cs.CV2025-03被引 1

用草图双向交互提升部分损坏物体修复的精度与结构一致性

Recovering Partially Corrupted Objects via Sketch-Guided Bidirectional Feature Interaction

  • 草图与未损坏区域双向特征交互,增强空间控制
  • 在两个新数据集上优于现有最优方法,结构恢复更准确
  • 适合需要精细结构修复的图像编辑场景

文本引导的扩散模型在物体修复中取得显著进展,通过文本提示提供高层语义指导。然而,在部分损坏物体场景下,其缺乏像素级空间控制能力,尤其当关键未损坏线索仍存在时。草图引导方法虽能改善结构控制,但多采用单向映射,忽略未遮挡区域的上下文信息,导致草图与真实内容不一致。为此,我们提出基于预训练Stable Diffusion的草图引导双向特征交互框架。该框架包含两个互补方向:上下文到草图,将未损坏区域的多尺度潜在表示传递至草图分支,生成适配可见上下文与去噪进度的视觉掩码;草图到修复,通过草图条件仿射变换调节草图引导强度,确保与未损坏内容一致。该交互机制在扩散U-Net编码器的多个尺度上应用,实现更高空间保真度的结构恢复。在两个新构建的基准数据集上的大量实验表明,本方法优于当前最优技术。

原文摘要 · Abstract (English)

Text-guided diffusion models have achieved remarkable success in object inpainting by providing high-level semantic guidance through text prompts. However, they often lack precise pixel-level spatial control, especially in scenarios involving partially corrupted objects where critical uncorrupted cues remain. To overcome this limitation, sketch-guided methods have been introduced, using either indirect gradient modulation or direct sketch injection to improve structural control. Yet, existing approaches typically establish a one-way mapping from the sketch to the masked regions only, neglecting the contextual information from unmasked object areas. This leads to a disconnection between the sketch and the uncorrupted content, thereby causing sketch-guided inconsistency and structural mismatch. To tackle this challenge, we propose a sketch-guided bidirectional feature interaction framework built upon a pretrained Stable Diffusion model. Our bidirectional interaction features two complementary directions, context-to-sketch and sketch-to-inpainting, that enable fine-grained spatial control for partially corrupted object inpainting. In the context-to-sketch direction, multi-scale latents from uncorrupted object regions are propagated to the sketch branch to generate a visual mask that adapts the sketch features to the visible context and denoising progress. In the sketch-to-inpainting direction, a sketch-conditional affine transformation modulates the influence of sketch guidance based on the learned visual mask, ensuring consistency with uncorrupted object content. This interaction is applied at multiple scales within the encoder of the diffusion U-Net, enabling the model to restore object structures with enhanced spatial fidelity. Extensive experiments on two newly constructed benchmark datasets demonstrate that our approach outperforms state-of-the-art methods.

图像修复草图引导扩散模型双向交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。