解决多层推理系统中反馈稀疏时的稳定路由学习问题
Online Learning for Multi-Layer Hierarchical Inference under Partial and Policy-Dependent Feedback
- 设计基于扩展指数梯度的低方差算法,结合李雅普诺夫优化
- 在反馈概率随深度递减的条件下实现无偏损失估计与稳定学习
- 适合资源受限、仅终端有反馈的实时推理系统部署
多层分层推理系统将任务在多个计算层间调度,每个节点可本地完成预测或向下一层次转发任务。在长期资源约束和仅终端反馈的条件下,学习最优调度策略面临挑战:预测误差仅在最终验证层暴露,导致反馈具有部分性与策略依赖性,观测概率随层级加深而衰减,使传统重要性加权方法方差剧增。本文形式化了递归损失结构,证明朴素的重要性加权上下文赌博机方法会因反馈概率下降而失稳。为此,提出一种结合李雅普诺夫优化的方差缩减型EXP4算法,实现无偏损失估计与稳定学习。理论证明该方法在随机到达与资源约束下接近最优,并提供相对于最优固定路由策略的后悔界。大规模多任务实验表明,相比标准重要性加权方法,本方法在稳定性与性能上均有显著提升。
原文摘要 · Abstract (English)
Hierarchical inference systems route tasks across multiple computational layers, where each node may either finalize a prediction locally or offload the task to a node in the next layer for further processing. Learning optimal routing policies in such systems is challenging: inference loss is defined recursively across layers, while feedback on prediction error is revealed only at a terminal oracle layer. This induces a partial, policy-dependent feedback structure in which observability probabilities decay with depth, causing importance-weighted estimators to suffer from amplified variance. We study online routing for multi-layer hierarchical inference under long-term resource constraints and terminal-only feedback. We formalize the recursive loss structure and show that naive importance-weighted contextual bandit methods become unstable as feedback probability decays along the hierarchy. To address this, we develop a variance-reduced EXP4-based algorithm integrated with Lyapunov optimization, yielding unbiased loss estimation and stable learning under sparse and policy-dependent feedback. We provide regret guarantees relative to the best fixed routing policy in hindsight and establish near-optimality under stochastic arrivals and resource constraints. Experiments on large-scale multi-task workloads demonstrate improved stability and performance compared to standard importance-weighted approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。