腾讯发布百万级广告生成推荐数据集,推动多模态推荐研究。
Tencent Advertising Algorithm Challenge 2025: All-Modality Generative Recommendation
- 构建真实工业场景下的多模态序列生成数据集
- 支持点击与转化事件区分,引入高价值目标加权评估
- 适合从事广告推荐、生成式AI研究者使用
生成式推荐系统正成为新范式,通过将协同标识与多模态内容映射到离散标记空间,并用自回归序列模型建模用户行为。尽管已有多种多模态推荐数据集,但缺乏公开的、大规模、真实且专为工业广告场景中的生成式推荐(GR)设计的基准数据。为此,我们组织了腾讯广告算法挑战赛2025,基于两个全模态数据集:TencentGR-1M 和 TencentGR-10M。两者均源自真实去标识化的腾讯广告日志,包含先进的嵌入模型提取的丰富协同标识与多模态表示。预赛阶段(TencentGR-1M)提供100万用户序列,每条序列最多包含100个交互项,每个交互标注曝光与点击信号;决赛阶段(TencentGR-10M)扩展至1000万用户,并在序列和目标层面明确区分点击与转化事件。本文介绍任务定义、数据构建流程、特征架构、基线GR模型、评估协议及顶尖方案的关键发现。我们的数据集聚焦广告场景下的多模态序列生成,并引入对高价值转化事件的加权评估。数据集已发布于 https://huggingface.co/datasets/TAAC2025,基线代码见 https://github.com/TencentAdvertisingAlgorithmCompetition/baseline_2025,官网为 https://algo.qq.com/2025。
原文摘要 · Abstract (English)
Generative recommender systems are rapidly emerging as a new paradigm for recommendation, where collaborative identifiers and/or multi-modal content are mapped into discrete token spaces and user behavior is modelled with autoregressive sequence models. Despite progress on multi-modal recommendation datasets, there is still a lack of public benchmarks that jointly offer large-scale, realistic and fully all-modality data designed specifically for generative recommendation (GR) in industrial advertising. To foster research in this direction, we organised the Tencent Advertising Algorithm Challenge 2025, a global competition built on top of two all-modality datasets for GR: TencentGR-1M and TencentGR-10M. Both datasets are constructed from real de-identified Tencent Ads logs and contain rich collaborative IDs and multi-modal representations extracted with state-of-the-art embedding models. The preliminary track (TencentGR-1M) provides 1 million user sequences with up to 100 interacted items each, where each interaction is labeled with exposure and click signals, while the final track (TencentGR-10M) scales this to 10 million users and explicitly distinguishes between click and conversion events at both the sequence and target level. This paper presents the task definition, data construction process, feature schema, baseline GR model, evaluation protocol, and key findings from top-ranked and award-winning solutions. Our datasets focus on multi-modal sequence generation in an advertising setting and introduce weighted evaluation for high-value conversion events. We release our datasets at https://huggingface.co/datasets/TAAC2025 and baseline implementations at https://github.com/TencentAdvertisingAlgorithmCompetition/baseline_2025 to enable future research on all-modality generative recommendation at an industrial scale. The official website is https://algo.qq.com/2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。