arXiv:2508.12628cs.CV2025-08被引 1

用大模型让广告图选择更智能可解释,还能懂用户喜好。

Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning

  • 用多模态大模型将选图转为自然语言推理任务,实现可解释评估。
  • 构建8000对标注的对比图像数据集,支持模型学习优劣判断。
  • 适合需要精准选图的电商广告主和平台,提升投放效果。

广告创意图是电商平台的核心,吸引眼球的图片能提升用户体验,增加广告商收益与平台收入。随着AIGC技术发展,广告商可低成本生成大量创意图,但难以评估质量进行筛选。现有方法多聚焦于排序,缺乏可解释性。本文提出首个可解释创意评估与选择范式,基于多模态大语言模型(MLLMs),将创意图评估与选择统一为自然语言生成任务。为此,我们构建了首个基于对比推理的创意数据集CreativePair,包含8000个标注图像对,每对均标有更优图像。同时提出Creative4U(发音同Creative for You),一个考虑用户兴趣的MLLMs驱动选图系统。通过基于思维链监督微调(CoT-SFT)和组相对策略优化(GRPO)的强化学习训练流程(Reason-to-Select RFT),Creative4U可准确评估并选出最优创意图。离线与在线实验均验证了方法的有效性。代码与数据集将公开,以推动学术与工业应用。

原文摘要 · Abstract (English)

Creative image in advertising is the heart and soul of e-commerce platform. An eye-catching creative image can enhance the shopping experience for users, boosting income for advertisers and advertising revenue for platforms. With the advent of AIGC technology, advertisers can produce large quantities of creative images at minimal cost. However, they struggle to assess the creative quality to select. Existing methods primarily focus on creative ranking, which fails to address the need for explainable creative selection. In this work, we propose the first paradigm for explainable creative assessment and selection. Powered by multimodal large language models (MLLMs), our approach integrates the assessment and selection of creative images into a natural language generation task. To facilitate this research, we construct CreativePair, the first comparative reasoning-induced creative dataset featuring 8k annotated image pairs, with each sample including a label indicating which image is superior. Additionally, we introduce Creative4U (pronounced Creative for You), a MLLMs-based creative selector that takes into account users' interests. Through Reason-to-Select RFT, which includes supervised fine-tuning with Chain-of-Thought (CoT-SFT) and Group Relative Policy Optimization (GRPO) based reinforcement learning, Creative4U is able to evaluate and select creative images accurately. Both offline and online experiments demonstrate the effectiveness of our approach. Our code and dataset will be made public to advance research and industrial applications.

广告生成多模态大模型可解释性选图系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。