arXiv:2506.17612cs.CV2025-06NeurIPS被引 40

用AI代理实现像专业摄影师一样智能修图,还能自由调节细节。

JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

  • 基于多模态大模型的AI修图代理,理解用户意图并自动调用200多个Lightroom工具。
  • 在真实用户编辑数据集上,像素级保真度比GPT-4o高60%,且控制更精细。
  • 适合想高效创作又不想学复杂软件的设计师、摄影爱好者和内容创作者。

照片修饰已成为现代视觉叙事的核心,使用户能够捕捉美学并表达创意。尽管Adobe Lightroom等专业工具功能强大,但需要大量专业知识和手动操作。现有AI方案虽可自动化,却普遍存在可调性差、泛化能力弱的问题,难以满足多样化和个性化的编辑需求。为此,我们提出JarvisArt,一个由多模态大语言模型驱动的智能代理,能理解用户意图,模拟专业艺术家的推理过程,并智能协调Lightroom中超过200种修饰工具。该代理采用两阶段训练:先通过思维链监督微调建立基础推理与工具使用能力,再通过面向修图的组相对策略优化(GRPO-R)提升决策与工具熟练度。我们还设计了Agent-to-Lightroom协议,实现与Lightroom的无缝集成。为评估性能,我们构建了基于真实用户编辑的MMArt-Bench基准。实验表明,JarvisArt具备友好的交互体验、优异的泛化能力及对全局与局部调整的细粒度控制,显著优于GPT-4o——在MMArt-Bench上平均像素级指标提升60%的同时,保持相近的指令遵循能力。

原文摘要 · Abstract (English)

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based solutions provide automation but often suffer from limited adjustability and poor generalization, failing to meet diverse and personalized editing needs. To bridge this gap, we introduce JarvisArt, a multi-modal large language model (MLLM)-driven agent that understands user intent, mimics the reasoning process of professional artists, and intelligently coordinates over 200 retouching tools within Lightroom. JarvisArt undergoes a two-stage training process: an initial Chain-of-Thought supervised fine-tuning to establish basic reasoning and tool-use skills, followed by Group Relative Policy Optimization for Retouching (GRPO-R) to further enhance its decision-making and tool proficiency. We also propose the Agent-to-Lightroom Protocol to facilitate seamless integration with Lightroom. To evaluate performance, we develop MMArt-Bench, a novel benchmark constructed from real-world user edits. JarvisArt demonstrates user-friendly interaction, superior generalization, and fine-grained control over both global and local adjustments, paving a new avenue for intelligent photo retouching. Notably, it outperforms GPT-4o with a 60% improvement in average pixel-level metrics on MMArt-Bench for content fidelity, while maintaining comparable instruction-following capabilities. Project Page: https://jarvisart.vercel.app/.

智能修图多模态大模型应用AI创作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。