arXiv:2508.21090cs.CV2025-08

解决零样本风格迁移中的注意力泄漏问题,提升图像语义对齐精度

Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment

  • 通过查询-查询对齐实现图像间精细空间语义映射
  • 在多个数据集上显著提升风格保真度,结构保持能力也领先
  • 适合关注风格迁移质量与语义一致性的视觉生成研究者

我们发现,使用大规模图像生成模型进行零样本外观迁移时面临一个关键挑战:注意力泄漏。该问题源于两幅图像间的语义映射被查询-键对齐所捕获。为解决此问题,我们提出Q-Align,利用查询-查询对齐来缓解注意力泄漏,改善零样本外观迁移中的语义对齐。Q-Align包含三大贡献:(1) 查询-查询对齐,促进两图像间复杂的空间语义映射;(2) 键-值重排,通过重新对齐增强特征对应关系;(3) 使用重排后的键和值进行注意力细化,以维持语义一致性。通过大量实验与分析验证,Q-Align在外观保真度上优于现有最先进方法,同时保持了竞争力的结构保留能力。

原文摘要 · Abstract (English)

We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic mapping between two images is captured by the Query-Key alignment. To tackle this issue, we introduce Q-Align, utilizing Query-Query alignment to mitigate attention leakage and improve the semantic alignment in zero-shot appearance transfer. Q-Align incorporates three core contributions: (1) Query-Query alignment, facilitating the sophisticated spatial semantic mapping between two images; (2) Key-Value rearrangement, enhancing feature correspondence through realignment; and (3) Attention refinement using rearranged keys and values to maintain semantic consistency. We validate the effectiveness of Q-Align through extensive experiments and analysis, and Q-Align outperforms state-of-the-art methods in appearance fidelity while maintaining competitive structure preservation.

风格迁移注意力机制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。