arXiv:2510.04034cs.CVcs.AI2025-10

通过优化注意力机制提升文本控图编辑的精度与一致性

Prompt-to-Prompt: Text-Based Image Editing Via Cross-Attention Mechanisms -- The Research of Hyperparameters and Novel Mechanisms to Enhance Existing Frameworks

  • 提出注意力重加权方法,增强文本提示对图像生成的精准控制
  • 构建CL P2P框架,解决循环不一致等现有编辑方法缺陷
  • 揭示超参数与注意力结构对生成质量的关键影响,适合图像编辑研究者

近期图像编辑技术从手动像素操作转向基于深度学习的方法,如稳定扩散模型,其利用交叉注意力机制实现文本驱动控制。这一转变虽简化了编辑流程,但导致结果不稳定,例如发色变化不一致。本文旨在通过探索和优化超参数,提升提示到提示(Prompt-to-Prompt)图像编辑框架的精确性与可靠性。我们系统研究了“词替换”方法,提出“注意力重加权”以增强适应性,并设计“CL P2P”框架以解决循环不一致等现有局限。本研究深化了对超参数设置与神经网络架构(特别是注意力机制)之间关系的理解,这些因素显著影响生成图像的构图与质量。

原文摘要 · Abstract (English)

Recent advances in image editing have shifted from manual pixel manipulation to employing deep learning methods like stable diffusion models, which now leverage cross-attention mechanisms for text-driven control. This transition has simplified the editing process but also introduced variability in results, such as inconsistent hair color changes. Our research aims to enhance the precision and reliability of prompt-to-prompt image editing frameworks by exploring and optimizing hyperparameters. We present a comprehensive study of the "word swap" method, develop an "attention re-weight method" for better adaptability, and propose the "CL P2P" framework to address existing limitations like cycle inconsistency. This work contributes to understanding and improving the interaction between hyperparameter settings and the architectural choices of neural network models, specifically their attention mechanisms, which significantly influence the composition and quality of the generated images.

图像编辑扩散模型注意力机制文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。