arXiv:2510.08532cs.CVcs.AI2025-10中稿 · CVPR被引 13

让文字编辑图像时可调节修改强度,从轻微到完全改变

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

  • 用一个数值控制编辑力度,实现渐变式修改
  • 无需针对每种属性单独训练,支持多种编辑类型
  • 通过合成数据集训练,提升编辑精度和一致性

基于文本指令的图像编辑提供了直观的图像操作方式,但仅依赖文字指令难以实现精细的编辑程度控制。本文提出Kontinuous Kontext,一种支持连续编辑强度控制的指令驱动图像编辑模型。该模型在现有先进图像编辑架构基础上,新增一个标量输入——编辑强度值,与指令结合使用,实现从无变化到完全重构的平滑过渡。为注入该标量信息,设计轻量级投影网络,将强度值与指令映射至模型调制空间的系数。训练数据通过现有生成模型合成多样化图像-指令-强度四元组,并经筛选确保质量与一致性。Kontinuous Kontext统一支持风格化、属性、材质、背景、形状等多种编辑操作的细粒度强度调控,无需特定属性的额外训练。

原文摘要 · Abstract (English)

Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language. Yet, relying solely on text instructions limits fine-grained control over the extent of edits. We introduce Kontinuous Kontext, an instruction-driven editing model that provides a new dimension of control over edit strength, enabling users to adjust edits gradually from no change to a fully realized result in a smooth and continuous manner. Kontinuous Kontext extends a state-of-the-art image editing model to accept an additional input, a scalar edit strength which is then paired with the edit instruction, enabling explicit control over the extent of the edit. To inject this scalar information, we train a lightweight projector network that maps the input scalar and the edit instruction to coefficients in the model's modulation space. For training our model, we synthesize a diverse dataset of image-edit-instruction-strength quadruplets using existing generative models, followed by a filtering stage to ensure quality and consistency. Kontinuous Kontext provides a unified approach for fine-grained control over edit strength for instruction driven editing from subtle to strong across diverse operations such as stylization, attribute, material, background, and shape changes, without requiring attribute-specific training.

图像编辑文本控制强度调节生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。