arXiv:2601.14711cs.AIcs.LG2026-01中稿 · The ACM Web Confer…被引 3

用大模型少样本推理+精细优化,提升广告竞价收益

DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs

  • 分两阶段:先用提示词生成策略,再通过反馈优化
  • 在真实与合成数据上均超越现有方法,提升广告主总收益
  • 适合缺乏历史数据的广告主,尤其适用于少样本场景

在人工智能生成竞价(AIGB)范式下,如何在预算约束下最大化广告主赢得展示位的累计价值是一个复杂挑战。广告主常有个性化目标但历史交互数据有限,导致传统强化学习方法在少样本场景下表现不佳。大语言模型(LLMs)凭借其上下文学习能力,可从少量数据中泛化,但缺乏精细化优化所需的数值精度。为此,我们提出GRPO-Adaptive,一种高效的LLM后训练策略,通过动态更新参考策略,同时提升推理与数值精度。基于此,我们进一步构建DARA,一个双阶段框架:第一阶段为少样本推理器,通过上下文提示生成初始策略;第二阶段为细粒度优化器,利用反馈驱动推理对策略进行精修。该分离设计使DARA兼具大模型的上下文学习优势与AIGB任务所需的精准适应性。在真实世界与合成数据环境上的大量实验表明,本方法在预算约束下持续优于现有基线,显著提升广告主累计价值。

原文摘要 · Abstract (English)

Optimizing the advertiser's cumulative value of winning impressions under budget constraints poses a complex challenge in online advertising, under the paradigm of AI-Generated Bidding (AIGB). Advertisers often have personalized objectives but limited historical interaction data, resulting in few-shot scenarios where traditional reinforcement learning (RL) methods struggle to perform effectively. Large Language Models (LLMs) offer a promising alternative for AIGB by leveraging their in-context learning capabilities to generalize from limited data. However, they lack the numerical precision required for fine-grained optimization. To address this limitation, we introduce GRPO-Adaptive, an efficient LLM post-training strategy that enhances both reasoning and numerical precision by dynamically updating the reference policy during training. Built upon this foundation, we further propose DARA, a novel dual-phase framework that decomposes the decision-making process into two stages: a few-shot reasoner that generates initial plans via in-context prompting, and a fine-grained optimizer that refines these plans using feedback-driven reasoning. This separation allows DARA to combine LLMs' in-context learning strengths with precise adaptability required by AIGB tasks. Extensive experiments on both real-world and synthetic data environments demonstrate that our approach consistently outperforms existing baselines in terms of cumulative advertiser value under budget constraints.

广告竞价大模型少样本学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。