改进电力拍卖学习率,实现更优竞标策略收敛速度。
Improved learning rates in multi-unit uniform price auctions
- 构建新竞标空间模型,利用问题结构优化学习算法。
- 带宽反馈下实现$ ilde{O}(K^{4/3}T^{2/3})$后悔率,优于此前$ ilde{O}(K^{7/4}T^{3/4})$。
- 引入胜标价全揭示反馈,适用于电力市场等实际场景。
受电力生产商参与日前电力市场策略行为的启发,本文研究重复多单位统一价格拍卖中的在线学习问题,聚焦对抗性对手出价设定。主要贡献在于提出一种新的竞标空间建模方法。证明了利用该问题结构的学习算法在带宽反馈下可达到$ ilde{O}(K^{4/3}T^{2/3})$的后悔率,优于文献中先前的$ ilde{O}(K^{7/4}T^{3/4})$。该结果在对数因子内为紧致。受电力备用市场启发,进一步引入一种新反馈机制:所有胜标价均被揭示。该反馈根据拍卖结果介于完全信息与带宽情形之间。证明所提出的算法在此反馈下可达到$ ilde{O}(K^{5/2} ootrom{T})$的后悔率。
原文摘要 · Abstract (English)
Motivated by the strategic participation of electricity producers in electricity day-ahead market, we study the problem of online learning in repeated multi-unit uniform price auctions focusing on the adversarial opposing bid setting. The main contribution of this paper is the introduction of a new modeling of the bid space. Indeed, we prove that a learning algorithm leveraging the structure of this problem achieves a regret of $\tilde{O}(K^{4/3}T^{2/3})$ under bandit feedback, improving over the bound of $\tilde{O}(K^{7/4}T^{3/4})$ previously obtained in the literature. This improved regret rate is tight up to logarithmic terms. Inspired by electricity reserve markets, we further introduce a different feedback model under which all winning bids are revealed. This feedback interpolates between the full-information and bandit scenarios depending on the auctions' results. We prove that, under this feedback, the algorithm that we propose achieves regret $\tilde{O}(K^{5/2}\sqrt{T})$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。