arXiv:2606.28943cs.CLcs.LG2026-06

A3M框架提升竞拍策略,自适应对抗并兼顾多方目标。

A3M: Adaptive, Adversarial and Multi-Objective Learning for Strategic Bidding in Repeated Auctions

论文配图:A3M: Adaptive, Adversarial and Multi-Objective Learning for Strategic Bidding in Repeated Auctions
图 1 · 摘自论文原文
  • 用深度强化学习动态平衡探索与利用,结合对手模型应对非平稳对手。
  • 相比基线减少30%~40%最终遗憾,且在多单位竞拍中表现稳定。
  • 适合需应对复杂竞拍环境、追求公平与收益平衡的研究者使用。

在具有弱反馈的重复多单位拍卖中学习竞拍策略面临根本性挑战。现有方法通常依赖固定的探索-利用调度,假设对手为静态,仅优化竞标方收益,限制了适应性与战略鲁棒性。为此,我们提出A3M框架,融合自适应深度强化学习(DRL)、显式对抗推理与合理多目标奖励设计,实现在线竞拍策略优化。A3M采用演员-评论家DRL主干,动态调节探索与利用;通过对手模型进行虚拟博弈以应对非平稳对手;设计复合奖励函数,联合最大化收益、拍卖方收入与公平性。我们在区分价格与统一价格拍卖中首次全面评估该方法,结果表明:在标准设置下,A3M使最终遗憾降低30%~40%,对对手策略突变保持鲁棒,随单位数K增长表现良好,并支持可调的多目标权衡。大量消融实验验证了各核心组件的必要性。本工作确立了A3M作为复杂拍卖环境中学习的强大灵活框架。

原文摘要 · Abstract (English)

Learning to bid in repeated multi-unit auctions with bandit feedback poses a fundamental challenge. Existing methods often rely on rigid explore-then-exploit schedules, assume stationary adversaries, and optimize solely for bidder utility, thereby limiting adaptability and strategic robustness. To address these limitations, we introduce the A3M framework, which integrates adaptive deep reinforcement learning (DRL), explicit adversarial reasoning, and principled multi-objective reward design for online auction strategy optimization. A3M employs an actor-critic DRL backbone to dynamically balance exploration and exploitation, an opponent model for fictitious play against non-stationary adversaries, and a composite reward function to jointly maximize utility, auctioneer revenue, and fairness. We provide the first comprehensive empirical evaluation of this integrated approach against established baselines in both discriminatory and uniform price auctions. Results show that A3M reduces final regret by 30--40\% in standard settings, maintains robust performance against adversarial strategy shifts, scales favorably with the number of units $K$, and enables tunable multi-objective trade-offs. An extensive ablation study confirms the necessity of each core component. Our work establishes A3M as a powerful and flexible framework for learning in complex auction environments.

竞拍学习强化学习多目标优化对抗策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。