arXiv:2511.23002cs.CV2025-11被引 26

自进化图像编辑代理,通过多模态思考和自我优化提升修图质量。

JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization

  • 采用多模态思维链,结合视觉与文本信息减少指令幻觉。
  • 在ArtEdit-Bench上比Nano-Banana提升18.95%,像素级保真度提升44.96%。
  • 适合需要高精度、自迭代能力的图像编辑场景与研究者。

基于智能体的编辑模型显著提升了交互体验、处理质量和创作灵活性。然而仍存在两大挑战:(1) 指令幻觉——仅依赖文本的思维链(CoT)推理因信息瓶颈难以避免事实错误;(2) 奖励劫持——动态策略优化对抗静态奖励模型,导致智能体利用奖励函数漏洞。为此,我们提出JarvisEvo,一种统一的图像编辑智能体,模拟专家设计师的迭代过程:编辑、选择工具、评估结果、反思决策以优化输出。其三大优势为:(1) 交错式多模态思维链(iMCoT)机制,增强指令遵循与编辑质量;(2) 编辑-评估协同优化(SEPO)框架,实现无需外部奖励的自进化,有效缓解奖励劫持;(3) 通过无缝集成Adobe Lightroom,支持全局与局部细粒度编辑。在ArtEdit-Bench上,JarvisEvo平均优于Nano-Banana 18.95%,像素级内容保真度提升44.96%。

原文摘要 · Abstract (English)

Agent-based editing models have substantially advanced interactive experiences, processing quality, and creative flexibility. However, two critical challenges persist: (1) instruction hallucination, text-only chain-of-thought (CoT) reasoning cannot fully prevent factual errors due to inherent information bottlenecks; (2) reward hacking, dynamic policy optimization against static reward models allows agents to exploit flaws in reward functions. To address these issues, we propose JarvisEvo, a unified image editing agent that emulates an expert human designer by iteratively editing, selecting appropriate tools, evaluating results, and reflecting on its own decisions to refine outcomes. JarvisEvo offers three key advantages: (1) an interleaved multimodal chain-of-thought (iMCoT) reasoning mechanism that enhances instruction following and editing quality; (2) a synergistic editor-evaluator policy optimization (SEPO) framework that enables self-improvement without external rewards, effectively mitigating reward hacking; and (3) support for both global and local fine-grained editing through seamless integration of Adobe Lightroom. On ArtEdit-Bench, JarvisEvo outperforms Nano-Banana by an average of 18.95% on preservative editing metrics, including a substantial 44.96% improvement in pixel-level content fidelity. Project page: https://jarvisevo.vercel.app/

图像编辑智能体自进化多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。