提出可保证的专家模型复用、新建或延迟决策机制,兼顾实时性与可靠性。
Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools

- 基于条件差异的单边序贯检验,实现三种决策的统计可解释性
- 在多数据流上零误新建、零误复用,性能媲美甚至超越传统启发式方法
- 适用于持续学习系统中长期运行的专家池管理,尤其适合高重访场景
维护专家模型池的流式系统需反复决定是否复用现有专家、新建新专家或推迟决策。本文提出一个决策层,使三种结果均具有统计意义。复用与新建被建模为条件(机制级)差异的单边序贯假设,以不确定区间分隔;推迟即双方赌态过程均未积累足够证据的状态。证明了可预测判别器序列可观测代理差异的有限时间任意时刻有效性,并实现了对总体量的单边无条件转移,每侧松弛量等于单个判别器的超额风险;经验观察到的向下偏差规律使新建侧恰好保守。通过重启的e检测器(几何间隔重启的无窗赌注超鞅银行,内存复杂度O(log t)),在重启实例间分配误差预算,保持全生命周期任意时刻有效性;在专家创建顺序上支出同样控制多重性,支持无限多个专家。在合成多概念流、Electricity、Covertype及高重复性的INSECTS基准上,实例计数重启银行在切换后实现零误新建与零误复用,且在INSECTS重访准确率0.675上匹配或超过退役窗口化启发式方法,使部署算法与保证算法一致。
原文摘要 · Abstract (English)
Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn a new one, or defer. We present a decision layer that makes all three outcomes statistically meaningful. Reuse and spawn are posed as one-sided sequential hypotheses on a conditional (mechanism-level) discrepancy, separated by an indifference zone; defer is exactly the state in which neither betting e-process has accumulated sufficient evidence. We prove finite-time anytime validity for the observable surrogate discrepancy of a predictable discriminator sequence, and an unconditional one-sided transfer to the population quantity in which each side's slack is the excess risk of a single discriminator; an empirically observed downward-bias regularity makes the spawn side exactly conservative. Recency without sacrificing the guarantee is obtained by a restarted e-detector: a bank of unwindowed betting supermartingales at geometrically spaced restart times (O(log t) memory), with the error budget spent over restart instances, which preserves lifetime anytime validity; spending over expert-creation order likewise controls multiplicity for unboundedly many experts. On synthetic multi-concept streams, Electricity, Covertype, and the recurrence-heavy INSECTS benchmark, the instance-accounted restarted bank achieves zero false spawns and zero false reuses after switches and matches or exceeds the retired windowed heuristic (INSECTS-reoccurring accuracy 0.675), making the deployed algorithm and the guaranteed algorithm one and the same.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。