arXiv:2602.08388cs.CV2026-02被引 1

用扩散变压器实现精准且带真实光影的图像几何编辑

Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers

  • 通过上下文生成与几何变换融合,实现物体位置、角度、大小的精确调整
  • 在12万张图像对上训练,显著提升光影和阴影的真实感
  • 适合需要高精度几何编辑与视觉真实的图像处理场景

扩散模型虽大幅提升了图像编辑能力,但在复杂场景中处理平移、旋转、缩放等几何变换仍存在挑战。现有方法主要受限于:(1) 物体几何编辑精度不足;(2) 光照与阴影建模不充分,导致结果失真。为此,我们提出GeoEdit框架,利用扩散变压器模块进行上下文生成,并集成几何变换以实现精准物体编辑。同时引入效应敏感注意力机制,增强复杂光照与阴影建模。为支持训练,构建了包含超过12万对高质量图像的RS-Objects数据集,使模型在学习几何编辑的同时生成逼真的光影效果。在公开基准上的大量实验表明,GeoEdit在视觉质量、几何准确性和真实性方面均优于现有最先进方法。

原文摘要 · Abstract (English)

Recent advances in diffusion models have significantly improved image editing. However, challenges persist in handling geometric transformations, such as translation, rotation, and scaling, particularly in complex scenes. Existing approaches suffer from two main limitations: (1) difficulty in achieving accurate geometric editing of object translation, rotation, and scaling; (2) inadequate modeling of intricate lighting and shadow effects, leading to unrealistic results. To address these issues, we propose GeoEdit, a framework that leverages in-context generation through a diffusion transformer module, which integrates geometric transformations for precise object edits. Moreover, we introduce Effects-Sensitive Attention, which enhances the modeling of intricate lighting and shadow effects for improved realism. To further support training, we construct RS-Objects, a large-scale geometric editing dataset containing over 120,000 high-quality image pairs, enabling the model to learn precise geometric editing while generating realistic lighting and shadows. Extensive experiments on public benchmarks demonstrate that GeoEdit consistently outperforms state-of-the-art methods in terms of visual quality, geometric accuracy, and realism.

图像编辑扩散模型几何变换光影建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。