arXiv:2504.01603cs.CV2025-04CVPR被引 3

让图像主体位置可变,实现更自然的背景补全。

A$^\text{T}$A: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting

  • 设计自适应位移模块,根据特征动态调整主体位置。
  • 采用逆向特征变换,从深层到浅层优化位置预测。
  • 支持位置可调或固定,兼容多种生成需求。

图像修复旨在填补图像缺失区域。近年来,前景条件下的背景修复成为研究热点,即在提供前景主体和文本提示的情况下修复背景。现有方法通常严格保持主体原始位置,导致主体与生成背景不协调。为此,本文提出新任务——文本引导的主体位置可变背景修复,并引入自适应变换代理(A$^\text{T}$A)解决该问题。首先,设计PosAgent模块,基于输入特征自适应预测合适位移,实现主体位置变化;其次,提出逆向位移变换(RDT)模块,以反向结构将深层到浅层的特征图按语义信息进行变换;最后,引入位置切换嵌入,控制主体位置是否自适应预测或固定。大量对比实验验证了A$^\text{T}$A的有效性,不仅在主体位置可变修复中表现优异,同时在固定位置修复任务上也保持良好性能。

原文摘要 · Abstract (English)

Image inpainting aims to fill the missing region of an image. Recently, there has been a surge of interest in foreground-conditioned background inpainting, a sub-task that fills the background of an image while the foreground subject and associated text prompt are provided. Existing background inpainting methods typically strictly preserve the subject's original position from the source image, resulting in inconsistencies between the subject and the generated background. To address this challenge, we propose a new task, the "Text-Guided Subject-Position Variable Background Inpainting", which aims to dynamically adjust the subject position to achieve a harmonious relationship between the subject and the inpainted background, and propose the Adaptive Transformation Agent (A$^\text{T}$A) for this task. Firstly, we design a PosAgent Block that adaptively predicts an appropriate displacement based on given features to achieve variable subject-position. Secondly, we design the Reverse Displacement Transform (RDT) module, which arranges multiple PosAgent blocks in a reverse structure, to transform hierarchical feature maps from deep to shallow based on semantic information. Thirdly, we equip A$^\text{T}$A with a Position Switch Embedding to control whether the subject's position in the generated image is adaptively predicted or fixed. Extensive comparative experiments validate the effectiveness of our A$^\text{T}$A approach, which not only demonstrates superior inpainting capabilities in subject-position variable inpainting, but also ensures good performance on subject-position fixed inpainting.

图像修复文本生成位置可变AI作画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。