arXiv:2411.01545cs.CV2024-11中稿 · ACMMM 2024被引 7

提出无需训练的小物体编辑方法,解决文本生成小物体精度低问题

Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach

  • 通过局部与全局注意力引导,实现无需训练的精准小物体编辑
  • 在SOEBench数据集上,小物体生成质量显著优于现有模型
  • 适合需要高精度图像生成的视觉设计、工业质检等场景

近年来,基于大规模扩散模型(如Stable Diffusion)的文本引导图像编辑方法取得了显著进展。尽管扩散模型能生成高质量图像,但在小物体生成方面受限于文本与物体间跨模态注意力对齐困难。本文提出一种无需训练的方法,通过局部和全局注意力引导,有效缓解该对齐问题,显著提升模型根据文本描述准确渲染小物体的能力。同时,我们构建了标准化基准数据集SOEBench(Small Object Editing),基于MSCOCO和OpenImage收集,用于定量评估文本驱动的小物体生成。初步实验表明,所提方法在生成保真度和准确性上均优于现有模型。该成果不仅推动人工智能与计算机视觉发展,也为需精确图像生成的行业应用开辟新路径。数据集将公开于项目主页:https://soebench.github.io/

原文摘要 · Abstract (English)

A plethora of text-guided image editing methods has recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models especially Stable Diffusion. Despite the success of diffusion models in producing high-quality images, their application to small object generation has been limited due to difficulties in aligning cross-modal attention maps between text and these objects. Our approach offers a training-free method that significantly mitigates this alignment issue with local and global attention guidance , enhancing the model's ability to accurately render small objects in accordance with textual descriptions. We detail the methodology in our approach, emphasizing its divergence from traditional generation techniques and highlighting its advantages. What's more important is that we also provide~\textit{SOEBench} (Small Object Editing), a standardized benchmark for quantitatively evaluating text-based small object generation collected from \textit{MSCOCO} and \textit{OpenImage}. Preliminary results demonstrate the effectiveness of our method, showing marked improvements in the fidelity and accuracy of small object generation compared to existing models. This advancement not only contributes to the field of AI and computer vision but also opens up new possibilities for applications in various industries where precise image generation is critical. We will release our dataset on our project page: \href{https://soebench.github.io/}{https://soebench.github.io/}.

图像编辑小物体生成扩散模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。