arXiv:2511.23105cs.CV2025-11被引 3

用数值精确控制图像编辑强度,让语言指令更精准。

NumeriKontrol: Adding Numeric Control to Diffusion Transformers for Instruction-based Image Editing

  • 通过数值适配器将数字参数注入扩散模型,实现连续调节。
  • 在多种属性编辑中实现高精度、稳定可控的尺度调节效果。
  • 适合需要精细调整图像的设计师或交互式编辑用户。

基于指令的图像编辑可通过自然语言命令实现直观操作,但仅靠文本指令往往难以实现对编辑强度的精细控制。我们提出NumeriKontrol框架,允许用户使用带有常见单位的连续标量值精确调节图像属性。该方法通过高效的数值适配器编码数值编辑尺度,并以即插即用方式注入扩散模型。得益于任务分离的设计,其支持零样本多条件编辑,可任意顺序指定多个指令。为提供高质量监督信号,我们从高保真渲染引擎和DSLR相机等可靠来源合成训练数据。所构建的通用属性变换(CAT)数据集涵盖多样化的属性操作,具备准确的真值尺度,使NumeriKontrol可作为简洁而强大的交互式编辑工具。大量实验表明,NumeriKontrol在多种属性编辑场景中均能实现准确、连续且稳定的尺度控制。该工作推动了基于指令的图像编辑向更精确、可扩展和用户可控的方向发展。

原文摘要 · Abstract (English)

Instruction-based image editing enables intuitive manipulation through natural language commands. However, text instructions alone often lack the precision required for fine-grained control over edit intensity. We introduce NumeriKontrol, a framework that allows users to precisely adjust image attributes using continuous scalar values with common units. NumeriKontrol encodes numeric editing scales via an effective Numeric Adapter and injects them into diffusion models in a plug-and-play manner. Thanks to a task-separated design, our approach supports zero-shot multi-condition editing, allowing users to specify multiple instructions in any order. To provide high-quality supervision, we synthesize precise training data from reliable sources, including high-fidelity rendering engines and DSLR cameras. Our Common Attribute Transform (CAT) dataset covers diverse attribute manipulations with accurate ground-truth scales, enabling NumeriKontrol to function as a simple yet powerful interactive editing studio. Extensive experiments show that NumeriKontrol delivers accurate, continuous, and stable scale control across a wide range of attribute editing scenarios. These contributions advance instruction-based image editing by enabling precise, scalable, and user-controllable image manipulation.

图像编辑扩散模型数值控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。