用语言模型模拟专业摄影师的修图思维,让普通人也能做出有艺术感的调色。
PhotoArtAgent: Intelligent Photo Retouching with Language Model-Based Artist Agents
- 用视觉语言模型分析照片,规划分步调色策略
- 通过API控制Lightroom精准调整参数,自动迭代优化效果
- 每一步都生成可读的解释,让用户看得懂、控得住
照片修图不仅是技术处理,更是情感表达与叙事深化的艺术过程。虽然专业艺术家能通过精心调整创造独特视觉效果,但普通用户依赖的自动化工具往往只追求美观,缺乏解释性与交互透明度。本文提出PhotoArtAgent,一个结合视觉语言模型(VLMs)与自然语言推理的智能系统,模拟专业艺术家的创作流程:先进行显式艺术分析,制定修图策略,再通过API向Lightroom输出精确参数;随后评估生成图像,迭代优化直至达成预期艺术愿景。整个过程提供透明的文本解释,增强用户理解和控制力。实验表明,PhotoArtAgent在用户研究中优于现有自动化工具,并达到与专业人类艺术家相当的效果。
原文摘要 · Abstract (English)
Photo retouching is integral to photographic art, extending far beyond simple technical fixes to heighten emotional expression and narrative depth. While artists leverage expertise to create unique visual effects through deliberate adjustments, non-professional users often rely on automated tools that produce visually pleasing results but lack interpretative depth and interactive transparency. In this paper, we introduce PhotoArtAgent, an intelligent system that combines Vision-Language Models (VLMs) with advanced natural language reasoning to emulate the creative process of a professional artist. The agent performs explicit artistic analysis, plans retouching strategies, and outputs precise parameters to Lightroom through an API. It then evaluates the resulting images and iteratively refines them until the desired artistic vision is achieved. Throughout this process, PhotoArtAgent provides transparent, text-based explanations of its creative rationale, fostering meaningful interaction and user control. Experimental results show that PhotoArtAgent not only surpasses existing automated tools in user studies but also achieves results comparable to those of professional human artists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。