arXiv:2602.02033cs.CVcs.AI2026-02被引 4

针对广告图像生成中用户偏好差异问题,提出可适配多群体偏好的统一框架。

One Size, Many Fits: Aligning Diverse Group-Wise Click Preferences in Large-Scale Advertising Image Generation

  • 按用户与产品特征动态分组,构建群体偏好特征。
  • 用群体条件生成模型提升各群体点击率,线上表现最优。
  • 首次公开60万组规模的广告图像偏好数据集,适合广告与推荐研究者。

广告图像生成越来越关注点击率(CTR)等在线指标,但现有方法采用“一刀切”策略,在优化整体CTR的同时忽视了用户群体间的偏好差异,导致特定群体表现不佳,影响精准营销效果。为此,本文提出统一框架 One Size, Many Fits(OSMF),实现大规模广告图像生成中的多群体点击偏好对齐。OSMF首先通过产品感知的自适应分组,基于用户属性与产品特征动态划分群体,并用丰富的集体偏好特征表征每个群体。在此基础上,采用群体感知的多模态大语言模型(G-MLLM)生成个性化图像,该模型在预训练阶段同时理解群体特征并生成广告图。随后,通过提出的群体强化学习算法Group-DPO对G-MLLM进行微调,有效提升各群体生成图像的点击率。为推动该领域发展,本文构建首个大规模公开的群体化广告图像偏好数据集GAIP,涵盖4000万用户形成的约60万组群体。大量实验表明,该框架在离线与在线设置下均达到当前最优性能。代码与数据将开源。

原文摘要 · Abstract (English)

Advertising image generation has increasingly focused on online metrics like Click-Through Rate (CTR), yet existing approaches adopt a ``one-size-fits-all" strategy that optimizes for overall CTR while neglecting preference diversity among user groups. This leads to suboptimal performance for specific groups, limiting targeted marketing effectiveness. To bridge this gap, we present \textit{One Size, Many Fits} (OSMF), a unified framework that aligns diverse group-wise click preferences in large-scale advertising image generation. OSMF begins with product-aware adaptive grouping, which dynamically organizes users based on their attributes and product characteristics, representing each group with rich collective preference features. Building on these groups, preference-conditioned image generation employs a Group-aware Multimodal Large Language Model (G-MLLM) to generate tailored images for each group. The G-MLLM is pre-trained to simultaneously comprehend group features and generate advertising images. Subsequently, we fine-tune the G-MLLM using our proposed Group-DPO for group-wise preference alignment, which effectively enhances each group's CTR on the generated images. To further advance this field, we introduce the Grouped Advertising Image Preference Dataset (GAIP), the first large-scale public dataset of group-wise image preferences, including around 600K groups built from 40M users. Extensive experiments demonstrate that our framework achieves the state-of-the-art performance in both offline and online settings. Our code and datasets will be released at https://github.com/JD-GenX/OSMF.

广告生成群体偏好多模态模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。