arXiv:2606.14929cs.LGcs.AI2026-06被引 1

为嵌入模型路由设计了能自适应低秩结构的高效在线学习算法。

Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts

  • 将路由问题建模为带低秩专家的对抗性上下文线性老虎机。
  • 提出HPG算法,实现$ ilde{/mathcal O}(s oot M T)$的线性化策略后悔值。
  • 适用于高维、模型有限且反馈不完整的推荐系统场景。

现代推荐系统越来越多地动态将不同查询路由到多个嵌入模型。尽管这一问题具有重要实践意义,但在对抗性查询、赌博式反馈和模型可观测性受限等现实条件下仍缺乏深入理解。本文将嵌入模型路由形式化为带有低秩专家的对抗性上下文线性老虎机,其中上下文为查询,动作为项目,专家为在低秩潜在表示空间中工作的嵌入模型。首先,我们证明标准后悔度量存在结构误设或统计不可行性,并识别出一个对数二次策略类,该类足够表达查询相关的模型路由,同时具备高效在线学习所需的结构。其次,我们提出一种策略梯度算法Hypentropy Policy Gradient(HPG),其在信息不完整的情况下可证明自适应未知低秩结构,达到$ ilde{ ext{O}}(s oot M T)$的线性化策略后悔值,其中$s$为专家固有秩,$M$为模型数量,$T$为轮次数,从而避免维度灾难。最后,我们还提供了一种计算高效且无需调参的HPG实现。

原文摘要 · Abstract (English)

Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly understood under realistic conditions like adversarial queries, bandit feedback, and limited observability of models. We formalize embedding model routing as an adversarial contextual linear bandit with low-rank experts, where contexts are queries, actions are items, and experts are the embedding models working on low-rank latent representation spaces. We first establish that standard regret notions suffer from structural misspecification or statistical intractability, and we identify a log-quadratic policy class that is expressive enough to capture query-dependent model routing, yet structured enough to allow efficient online learning. Second, we propose a policy gradient algorithm called Hypentropy Policy Gradient (HPG). It provably adapts to the unknown low-rank structure under incomplete information and attains $\tilde{\mathcal O}(s\sqrt{M T})$ linearized policy regret -- where $s, M$, and $T$ are the intrinsic rank of the experts, the number of models, and the number of rounds -- thus avoiding a curse of dimensionality. Finally, we also provide an computationally efficient and parameter-free implementation of HPG.

模型路由在线学习低秩结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。