arXiv:2602.17558cs.CV2026-02被引 2

用通用奖励模型让AI理解指令并自动修图,效果更自然。

RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

  • 用多模态大模型生成可执行的修图指令,连接用户意图与参数控制。
  • 通过自动生成评估指标,实现对主观修图效果的高质量反馈。
  • 适合需要智能、可解释修图助手的设计师和内容创作者。

近年来,多模态大语言模型(MLLM)在拓展视觉-语言推理至专业工具化图像编辑方面展现出巨大潜力,支持直观且富有创造力的编辑。一个有前景的方向是利用强化学习(RL)使MLLM能够推理并执行专业图像编辑软件中的最优工具使用计划。然而,由于缺乏可靠、可验证的奖励信号来反映创意编辑的主观性,训练仍具挑战。本文提出RetouchIQ框架,通过由通用奖励模型指导的MLLM代理,实现基于指令的可执行图像编辑。RetouchIQ能解析用户指定的编辑意图,并生成相应的可执行调整,将高层次审美目标与精确参数控制相衔接。为超越传统依赖固定参考图像的规则式奖励,我们提出一种通用奖励模型——经强化学习微调的MLLM,能针对每例生成相应评估指标进行判别。该模型通过多模态推理提供标量反馈,支持高质量、与指令一致的梯度更新。我们构建了一个包含19万条指令-推理对的扩展数据集,并建立新的基准用于指令式图像编辑。实验表明,RetouchIQ在语义一致性和感知质量上显著优于先前的MLLM及扩散模型编辑系统。结果表明,通用奖励驱动的MLLM代理在专业图像编辑中具有作为灵活、可解释、可执行助手的巨大潜力。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling intuitive and creative editing. A promising direction is to use reinforcement learning (RL) to enable MLLMs to reason about and execute optimal tool-use plans within professional image-editing software. However, training remains challenging due to the lack of reliable, verifiable reward signals that can reflect the inherently subjective nature of creative editing. In this work, we introduce RetouchIQ, a framework that performs instruction-based executable image editing through MLLM agents guided by a generalist reward model. RetouchIQ interprets user-specified editing intentions and generates corresponding, executable image adjustments, bridging high-level aesthetic goals with precise parameter control. To move beyond conventional, rule-based rewards that compute similarity against a fixed reference image using handcrafted metrics, we propose a generalist reward model, an RL fine-tuned MLLM that evaluates edited results through a set of generated metrics on a case-by-case basis. Then, the reward model provides scalar feedback through multimodal reasoning, enabling reinforcement learning with high-quality, instruction-consistent gradients. We curate an extended dataset with 190k instruction-reasoning pairs and establish a new benchmark for instruction-based image editing. Experiments show that RetouchIQ substantially improves both semantic consistency and perceptual quality over previous MLLM-based and diffusion-based editing systems. Our findings demonstrate the potential of generalist reward-driven MLLM agents as flexible, explainable, and executable assistants for professional image editing.

图像编辑多模态强化学习生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。