提出新算法,让专家排名更真实且自适应方差变化。
Rank-Induced PL Mirror Descent: A Rank-Faithful Second-Order Algorithm for Sleeping Experts
- 直接在排名诱导的普拉克特-卢瑟模型参数空间更新
- 首次实现睡眠专家场景下既忠实于排名又自适应方差
- 适合关注排名公平性与动态调整的在线学习研究者
我们提出一种新算法——秩诱导普拉克特-卢瑟镜面下降(RIPLM),利用文献[2022]建立的秩基准与分布基准之间的结构等价性。不同于以往基于专家身份操作的方法,RIPLM 在秩诱导的普拉克特-卢瑟(PL)参数化空间中直接更新,确保每轮策略始终属于秩诱导分布类,从而保持与秩基准的等价性。据我们所知,RIPLM 是首个在睡眠专家设置中同时具备(i)秩忠实性与(ii)方差自适应性的算法。
原文摘要 · Abstract (English)
We introduce a new algorithm, \emph{Rank-Induced Plackett--Luce Mirror Descent (RIPLM)}, which leverages the structural equivalence between the \emph{rank benchmark} and the \emph{distributional benchmark} established in \citet{BergamOzcanHsu2022}. Unlike prior approaches that operate on expert identities, RIPLM updates directly in the \emph{rank-induced Plackett--Luce (PL)} parameterization. This ensures that the algorithm's played distributions remain within the class of rank-induced distributions at every round, preserving the equivalence with the rank benchmark. To our knowledge, RIPLM is the first algorithm that is both (i) \emph{rank-faithful} and (ii) \emph{variance-adaptive} in the sleeping experts setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。