arXiv:2607.19344cs.CVcs.AI2026-07

用视觉提示精准控制图像生成中不同区域的外观。

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

论文配图:Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
图 1 · 摘自论文原文
  • 引入外观指针,让扩散模型按需定位文本或图像提示的位置。
  • 单模型在多项指标上超越专用模型,且不增加额外参数负担。
  • 适合需要精细局部控制的创意设计、图像编辑场景。

可控图像生成对创意工作者仍具挑战性,他们常需对材质、物体身份和空间布局进行精确的区域控制,仅靠文本提示难以实现。扩散变换器(DiTs)能原生处理来自文本和图像的异构标记,但缺乏决定这些标记应影响输出何处与如何影响的机制。本文提出外观指针——紧凑的标记,通过将文本或图像输入与用户指定的掩码对齐,引导DiT在正确空间位置获取正确的外观线索。外观指针由区域对应网络生成,并通过空间聚合机制优化,使模型在不显著增加标记负载的情况下处理多个区域描述。该方法首次在不重新训练基础模型的前提下,实现了模态无关的局部多模态控制。在多种评估指标上,单一模型性能达到或超过特定模态的最先进方法,为生成图像合成中的精确、区域感知、多模态引导提供了一条简单且可扩展的路径。

原文摘要 · Abstract (English)

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regional descriptions without significantly increasing token load. Our approach introduces the first modality-agnostic interface for localized multimodal control in a DiT without retraining the base model from scratch. Across a range of metrics, our single model reaches or surpasses the performance of modality-specific state of the art methods, offering a simple and extensible path toward precise, region-aware, multimodal guidance in generative image synthesis.

扩散模型图像生成区域控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。