分析8.3万真实修图需求,发现AI在精确修图上表现差于创意任务
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
- 基于Reddit 12年修图请求数据,研究真实用户编辑偏好
- 当前最佳AI(GPT-4o等)仅能完成约33%的请求,精确类任务更难
- AI常误改人/动物形象,且添加非请求修饰,适合创意型任务
生成式人工智能(GenAI)在自动化日常图像编辑任务方面具有巨大潜力,尤其自2025年3月25日GPT-4o发布以来。然而,人们最常希望修改什么内容?他们想进行哪些编辑操作(如移除或风格化主体)?是偏好结果可预测的精准编辑,还是高度创意的修改?通过分析过去12年(2013–2025)来自Reddit社区的8.3万条真实修图请求及对应的30.5万次专业修图师(PSR-wizard)编辑,我们回答了这些问题。根据人类评分,目前最优的AI编辑器(包括GPT-4o、Gemini-2.0-Flash、SeedEdit)仅能成功处理约33%的请求。有趣的是,AI在低创造性、需高精度的任务中表现反而劣于开放性任务。它们常无法保持人物和动物的身份一致性,且频繁添加未要求的修饰。另一方面,视觉语言模型(VLM)裁判(如o1)的判断与人类不同,可能更倾向AI生成的修改。代码与案例见:https://psrdataset.github.io
原文摘要 · Abstract (English)
Generative AI (GenAI) holds significant promise for automating everyday image editing tasks, especially following the recent release of GPT-4o on March 25, 2025. However, what subjects do people most often want edited? What kinds of editing actions do they want to perform (e.g., removing or stylizing the subject)? Do people prefer precise edits with predictable outcomes or highly creative ones? By understanding the characteristics of real-world requests and the corresponding edits made by freelance photo-editing wizards, can we draw lessons for improving AI-based editors and determine which types of requests can currently be handled successfully by AI editors? In this paper, we present a unique study addressing these questions by analyzing 83k requests from the past 12 years (2013-2025) on the Reddit community, which collected 305k PSR-wizard edits. According to human ratings, approximately only 33% of requests can be fulfilled by the best AI editors (including GPT-4o, Gemini-2.0-Flash, SeedEdit). Interestingly, AI editors perform worse on low-creativity requests that require precise editing than on more open-ended tasks. They often struggle to preserve the identity of people and animals, and frequently make non-requested touch-ups. On the other side of the table, VLM judges (e.g., o1) perform differently from human judges and may prefer AI edits more than human edits. Code and qualitative examples are available at: https://psrdataset.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。