无需微调的扩散模型实现精准区域编辑,支持全景图场景
Toward 360-Degree Indoor Panorama Editing via Tuning-Free Diffusion Model with Refocusing Cross-Attention

- 通过注意力重聚焦机制,只关注目标区域进行编辑
- 在30张多物体图像上实现94.3%的编辑精度与96.1%背景保真度
- 适合虚拟现实中的全景图像编辑,操作简单无需复杂提示
零样本文本引导的扩散模型虽推动了图像编辑发展,但仍面临三大挑战:提示词敏感需精细设计、编辑溢出影响非目标区域、小物体或杂乱场景因训练数据细粒度不足而失败。本文提出FocusDiff(目标感知注意力重聚焦),一种无需微调的精准区域编辑框架。基于自动分割或手动选择的目标区域,该方法对非编辑区域施加选择性模糊,引导注意力聚焦于目标区,同时准确传递对象的身份、结构和外观。集成上下文保持模块进一步保障背景一致性和全局连贯性,仅需单次运行即可实现由简单文本提示驱动的高质量编辑。我们还将FocusDiff拓展至360°室内全景图编辑,在虚拟现实环境中验证其有效性。在包含30张多物体图像及100个标注样本(含小物体案例)的局部编辑基准LIMB上,FocusDiff在文本-图像对齐与背景保留方面优于现有零样本编辑器,达到94.3%的编辑精度和96.1%的背景保真度,显著提升精确性、真实感与可用性。
原文摘要 · Abstract (English)
Zero-shot text-guided diffusion has significantly advanced image editing; however, its practical usability remains constrained by three persistent challenges: prompt brittleness that requires meticulous prompt engineering, spillover edits that unintentionally affect non-target regions, and failures on small or cluttered objects caused by limited fine-grained supervision in training data. We propose FocusDiff (Target-Aware Refocusing for Tuning-Free Diffusion Editing), a tuning-free framework for precise and region-specific image manipulation based on refocusing cross-attention. Given a target region obtained through automated segmentation or manual selection, FocusDiff applies selective blurring to non-edit areas to guide attention toward the masked region while accurately transferring the object's identity, structure, and appearance to the edited output. Integrated context-preserving modules further ensure background fidelity and global coherence, enabling accurate edits from simple text prompts in a single pass. We also extend FocusDiff to 360-degree indoor panorama editing and demonstrate its effectiveness within virtual reality environments. Extensive experiments on our localized editing benchmark LIMB, comprising 30 multi-object images and 100 annotated examples including challenging small-object cases, show that FocusDiff outperforms existing zero-shot editors in text-image alignment and background preservation, achieving superior precision, photorealism, and usability. The project page is available at https://vdkhoi20.github.io/FocusDiff.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。