arXiv:2604.25477cs.CVcs.AI2026-04被引 1

用双原子强化学习提升图像编辑的推理规划能力。

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

论文配图:DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing
图 1 · 摘自论文原文
  • 分离规划模块与生成模型,独立优化思考过程。
  • 通过认知与视觉双重原子奖励提升规划质量。
  • 适合研究图像编辑推理机制或想改进提示策略的人。

近期图像编辑模型虽具备高视觉保真度,但在需要复杂推理的任务上表现不足。为此,我们提出DDA-Thinker,一种以思考者(Thinker)为中心的框架,可在固定生成模型(Editor)下独立优化规划模块。该解耦范式便于对规划模块进行受控分析,并更清晰评估其在固定编辑器下的贡献。为有效引导思考者,我们引入双原子强化学习框架,将反馈分解为两个可验证检查清单:认知原子奖励用于直接评估思考者可执行计划的质量,视觉原子奖励用于评估最终图像质量。为提升检查清单质量,合成过程不仅基于源图和用户指令,还融合理想编辑场景的理性参考描述。此外,我们设计了两阶段数据构建流程:先生成多样化且聚焦推理的数据集,再通过难度感知精炼形成有效的强化学习训练课程。在RISE-Bench和KRIS-Bench等推理驱动图像编辑基准上的大量实验表明,该方法显著提升整体性能。所提方法使开源模型达到与强专有模型相当的效果,凸显了固定编辑器下以思考者为中心优化的实用潜力。

原文摘要 · Abstract (English)

Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance the reasoning-grounded planning for image editing, we propose DDA-Thinker, a Thinker-centric framework designed for the independent optimization of a planning module (Thinker) over a fixed generative model (Editor). This decoupled Thinker-centric paradigm facilitates a controlled analysis of the planning module and makes its contribution under a fixed Editor easier to assess. To effectively guide this Thinker, we introduce a dual-atomic reinforcement learning framework. This framework decomposes feedback into two distinct atomic rewards implemented through verifiable checklists: a cognitive-atomic reward to directly assess the quality of the Thinker's executable plan, which serves as the actionable outcome of the Thinker's reasoning, and a visual-atomic reward to assess the final image quality. To improve checklist quality, our checklist synthesis is grounded not only in the source image and user instruction but also in a rational reference description of the ideal post-edit scene. To support this training, we further develop a two-stage data curation pipeline that first synthesizes a diverse and reasoning-focused dataset, then applies difficulty-aware refinement to curate an effective training curriculum for reinforcement learning. Extensive experiments on reasoning-driven image editing benchmarks, including RISE-Bench and KRIS-Bench, demonstrate that our approach substantially improves overall performance. Our method enables a community model to achieve results competitive with strong proprietary models, highlighting the practical potential of Thinker-centric optimization under a fixed-editor setting.

图像编辑强化学习推理规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。