让内容推广兼顾短期曝光与长期模型优化,避免优质内容被误伤。
Guiding the Recommender: Information-Aware Auto-Bidding for Content Promotion
- 设计可分解的梯度覆盖目标,实时估算内容学习信号。
- 两阶段自动出价算法实现预算动态分配与每展现竞价优化。
- 适用于需要长期推荐效果提升的内容平台,尤其适合数据稀疏场景。
现代内容平台通过拍卖机制进行付费推广以缓解冷启动问题。我们的实证分析发现,该模式存在反直觉缺陷:推广虽能拯救低至中等质量内容,却可能因强制向次优受众曝光,污染互动信号,从而损害高质量内容的未来推荐表现。为此,我们将内容推广重构为兼顾短期价值获取与长期模型优化的双目标问题。提出可分解的代理目标——梯度覆盖,并建立其与Fisher信息及最优实验设计的理论关联。设计基于拉格朗日对偶的两阶段自动出价算法,通过影子价格动态调控预算,利用每展示边际效用优化出价。针对出价时标签缺失问题,提出置信度门控梯度启发式方法,以及适用于黑盒模型的零阶变体,可实时可靠估计学习信号。提供理论保证:复合目标具有单调次模性,在线拍卖中具亚线性后悔,且预算可行。在合成与真实数据集上的大量离线实验验证了该框架的有效性:相比基线,显著提升最终AUC/LogLoss,精准控制预算,且在梯度近似为零阶时仍保持性能。结果表明,战略性、信息感知的推广可有效提升长期模型表现与自然推荐效果,超越简单的曝光最大化策略。
原文摘要 · Abstract (English)
Modern content platforms offer paid promotion to mitigate cold start by allocating exposure via auctions. Our empirical analysis reveals a counterintuitive flaw in this paradigm: while promotion rescues low-to-medium quality content, it can harm high-quality content by forcing exposure to suboptimal audiences, polluting engagement signals and downgrading future recommendation. We recast content promotion as a dual-objective optimization that balances short-term value acquisition with long-term model improvement. To make this tractable at bid time in content promotion, we introduce a decomposable surrogate objective, gradient coverage, and establish its formal connection to Fisher Information and optimal experimental design. We design a two-stage auto-bidding algorithm based on Lagrange duality that dynamically paces budget through a shadow price and optimizes impression-level bids using per-impression marginal utilities. To address missing labels at bid time, we propose a confidence-gated gradient heuristic, paired with a zeroth-order variant for black-box models that reliably estimates learning signals in real time. We provide theoretical guarantees, proving monotone submodularity of the composite objective, sublinear regret in online auction, and budget feasibility. Extensive offline experiments on synthetic and real-world datasets validate the framework: it outperforms baselines, achieves superior final AUC/LogLoss, adheres closely to budget targets, and remains effective when gradients are approximated zeroth-order. These results show that strategic, information-aware promotion can improve long-term model performance and organic outcomes beyond naive impression-maximization strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。