arXiv:2607.17281cs.LGcs.AI2026-07被引 1

用分层推理框架让广告竞价自动进化,提升策略探索能力。

AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization

论文配图:AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization
图 1 · 摘自论文原文
  • 分层设计规划器与执行器,实现宏观策略与微观决策协同。
  • 在真实数据集上实现比基线高出18.3%的广告转化率优化。
  • 适合研究广告智能竞价、大模型应用落地的从业者参考。

自动竞价在在线广告中至关重要,可自动调整出价以优化广告主的商业目标。新兴的AI生成竞价(AIGB)范式广泛采用生成建模来优化出价策略,但受限于离线数据集的模式覆盖不足和任务状态理解能力弱,难以有效探索最优策略。大型语言模型(LLMs)凭借先验世界知识和推理能力,为克服这些限制提供了可能。然而,直接将LLMs应用于自动竞价面临数值精度有限、幻觉问题和推理延迟等挑战。为此,我们提出AIGB-R1,一种分层自演化自动竞价框架,旨在通过LLM的推理能力增强AI生成竞价。该框架包含高层规划模块用于宏观策略规划,以及低层执行模块用于细粒度决策制定。在此基础上,我们设计了基于经验的自演化循环,实现从累积经验中自主探索与优化策略。采用离线预训练与后训练对齐的两阶段流程,并构建交互式竞价模拟环境用于策略部署。此外,提出解耦组相对策略优化(D-GRPO),通过优势解耦实现端到端优化。在大规模公开数据集上的实验结果验证了AIGB-R1的有效性。

原文摘要 · Abstract (English)

Auto-bidding plays an essential role in online advertising, automatically adjusting bids for advertisers to optimize their commercial goals. The emerging AI-Generated Bidding (AIGB) paradigm widely adopts generative modeling to optimize bidding strategies, yet suffers from the limited mode coverage of offline datasets and inadequate task-state understanding, hindering effective exploration of optimal strategies. Large Language Models (LLMs), with prior world knowledge and reasoning capabilities, offer a promising approach to overcome these limitations. However, directly applying LLMs to auto-bidding tasks faces inherent challenges in limited numerical precision, hallucinations, and inference latency. To address these limitations, we propose AIGB-R1, a hierarchical self-evolving auto-bidding framework aiming to enhance AI-Generated Bidding via LLMs' Reasoning capabilities, comprising a high-level Planner module for macro-level strategy planning and a low-level Executor module for fine-grained decision-making. Building upon this, we design an experience-driven self-evolving loop, enabling autonomous strategy exploration and optimization from accumulated experience. We adopt a two-stage pipeline of offline pre-training and post-training alignment, and build an interactive bidding simulation environment for strategy rollout. Furthermore, we propose Decoupled Group Relative Policy Optimization (D-GRPO) to achieve end-to-end optimization via advantage decoupling. Experimental results on a large-scale public dataset demonstrate the effectiveness of AIGB-R1.

自动竞价大模型应用策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。