arXiv:2504.12833cs.CVcs.AI2025-04ICCV被引 2

用少量参考图和AI反馈,让图像编辑模型更懂指令、保持结构一致。

SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback

  • 通过视觉提示+在线强化学习,无需大量人工标注即可对齐用户指令。
  • 仅需5张参考图、10次训练即可完成复杂场景的精细编辑。
  • 适合需要精准可控图像生成的科研与机器人仿真应用。

本文提出SPIE:一种基于指令的图像编辑扩散模型的语义与结构后训练新方法,解决与用户指令对齐及输入图像一致性难题。我们设计了一种在线强化学习框架,通过人工智能反馈实现与人类偏好的对齐,无需依赖大量人工标注或大规模数据集构建。该方法在两个方面显著提升性能:首先,利用视觉提示捕捉期望编辑中的细微差别,实现无需冗长文本描述的精细化控制;其次,在复杂场景中实现精确且结构一致的修改,同时保持指令无关区域的高保真度。该方法极大简化了用户操作,仅需5张描绘特定概念的参考图像即可完成训练。实验表明,经过10次训练迭代,SPIE可在复杂场景中完成精细编辑。最后,我们将该方法应用于机器人领域,通过针对性图像编辑提升模拟环境的视觉真实感,使其更适合作为现实世界场景的替代测试平台。

原文摘要 · Abstract (English)

This paper presents SPIE: a novel approach for semantic and structural post-training of instruction-based image editing diffusion models, addressing key challenges in alignment with user prompts and consistency with input images. We introduce an online reinforcement learning framework that aligns the diffusion model with human preferences without relying on extensive human annotations or curating a large dataset. Our method significantly improves the alignment with instructions and realism in two ways. First, SPIE captures fine nuances in the desired edit by leveraging a visual prompt, enabling detailed control over visual edits without lengthy textual prompts. Second, it achieves precise and structurally coherent modifications in complex scenes while maintaining high fidelity in instruction-irrelevant areas. This approach simplifies users' efforts to achieve highly specific edits, requiring only 5 reference images depicting a certain concept for training. Experimental results demonstrate that SPIE can perform intricate edits in complex scenes, after just 10 training steps. Finally, we showcase the versatility of our method by applying it to robotics, where targeted image edits enhance the visual realism of simulated environments, which improves their utility as proxy for real-world settings.

图像编辑扩散模型强化学习机器人仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。