arXiv:2509.24514cs.CV2025-09

让图像编辑精准控制物体数量和位置,支持复杂场景多对象操作

Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency

  • 通过融合布局先验与视觉特征,增强空间结构理解
  • 在多物体场景中实现数量与布局一致性,性能超越现有模型
  • 适用于需要精确控制物体数量和位置的图像编辑任务

基于标准CLIP文本编码器的指令驱动图像编辑在多物体复杂场景中常失效。本文提出QL-Adapter框架,解决两个挑战:保持物体数量与空间布局一致性,并兼容多样类别。该框架包含两个核心模块:图像-布局融合模块(ILFM)将布局先验与CLIP图像编码器的ViT patch tokens融合,强化空间结构理解;跨模态增强模块(CMAM)将图像特征注入文本分支,丰富文本嵌入,提升指令遵循能力。我们还构建了QL-Dataset,一个涵盖广泛类别、布局与数量变化的基准数据集,并定义了数量与布局一致的图像编辑任务(QL-Edit)。大量实验表明,QL-Adapter在QL-Edit任务上达到当前最优性能,显著优于现有模型。

原文摘要 · Abstract (English)

Instruction driven image editing with standard CLIP text encoders often fails in complex scenes with many objects. We present QL-Adapter, a framework for multiple object editing that tackles two challenges: enforcing object counts and spatial layouts, and accommodating diverse categories. QL-Adapter consists of two core modules: the Image-Layout Fusion Module (ILFM) and the Cross-Modal Augmentation Module (CMAM). ILFM fuses layout priors with ViT patch tokens from the CLIP image encoder to strengthen spatial structure understanding. CMAM injects image features into the text branch to enrich textual embeddings and improve instruction following. We further build QL-Dataset, a benchmark that spans broad category, layout, and count variations, and define the task of quantity and layout consistent image editing (QL-Edit). Extensive experiments show that QL-Adapter achieves state of the art performance on QL-Edit and significantly outperforms existing models.

图像编辑多对象布局控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。