arXiv:2511.09715cs.CV2025-11被引 10

让图像编辑指令可滑动调节强度,实现精准连续控制。

SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control

  • 将多指令拆解为独立滑块,支持逐条精细调节
  • 仅用一组低秩适配矩阵,通用适配多种编辑任务
  • 适合需要交互式精细调整的图像编辑场景

基于指令的图像编辑模型虽已取得显著进展,能从多指令提示中完成复杂编辑,但各指令均以固定强度执行,难以精确控制编辑强度。本文提出SliderEdit框架,实现细粒度、可解释的连续指令控制。给定多部分编辑指令时,SliderEdit将各指令解耦并暴露为全局训练的滑块,支持平滑调节其强度。不同于以往需为每个属性单独训练或微调的滑块方法,本方法仅学习一组通用低秩适配矩阵,适用于多样编辑、属性及组合指令,实现沿各编辑维度的连续插值,同时保持空间局部性与全局语义一致性。我们在FLUX-Kontext和Qwen-Image-Edit等先进模型上应用SliderEdit,显著提升编辑可控性、视觉一致性和用户可操控性。据我们所知,这是首个探索并提出指令式图像编辑中连续细粒度控制的框架,为交互式、指令驱动的图像操作开辟新路径。

原文摘要 · Abstract (English)

Instruction-based image editing models have recently achieved impressive performance, enabling complex edits to an input image from a multi-instruction prompt. However, these models apply each instruction in the prompt with a fixed strength, limiting the user's ability to precisely and continuously control the intensity of individual edits. We introduce SliderEdit, a framework for continuous image editing with fine-grained, interpretable instruction control. Given a multi-part edit instruction, SliderEdit disentangles the individual instructions and exposes each as a globally trained slider, allowing smooth adjustment of its strength. Unlike prior works that introduced slider-based attribute controls in text-to-image generation, typically requiring separate training or fine-tuning for each attribute or concept, our method learns a single set of low-rank adaptation matrices that generalize across diverse edits, attributes, and compositional instructions. This enables continuous interpolation along individual edit dimensions while preserving both spatial locality and global semantic consistency. We apply SliderEdit to state-of-the-art image editing models, including FLUX-Kontext and Qwen-Image-Edit, and observe substantial improvements in edit controllability, visual consistency, and user steerability. To the best of our knowledge, we are the first to explore and propose a framework for continuous, fine-grained instruction control in instruction-based image editing models. Our results pave the way for interactive, instruction-driven image manipulation with continuous and compositional control.

图像编辑指令控制连续调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。