arXiv:2605.29139stat.MLcs.LG2026-05被引 1

让弱模型群在任意时间点都可靠地保证推理准确性,且通信成本更低。

Anytime-Valid Federated Conformal RAG for LLM Swarms

  • 用可累加的校准偏差预算,将固定时间覆盖转为任意时刻有效覆盖
  • 实验显示在多个数据集上实现14%-57%通信量节省,报警准确率不变
  • 适合需要动态调整资源的分布式语言模型系统

联邦共形检索增强生成(FC-RAG)在带宽受限的弱语言模型集群中提供分布无关的覆盖率,但仅限于固定时间点。本文将其扩展为任意时间有效的序列覆盖:在任意停止时刻均保持有效性,且在可预测的自适应控制(重校准、节点带宽提升、学生模型刷新)下仍成立,无需额外假设。传统组合方法失效,因FC-RAG的边际覆盖界导致投注过程的e-过程在不利校准下非超鞅,无法使用Ville不等式。为此提出Anytime-FC-RAG,基于可累加的每步校准偏差预算,将边际界转化为校准良好事件下的严格条件界,并采用截断的投注e-过程,使其在整个概率空间上为非负超鞅。由此获得四项保障:时间一致的报警有效性 $\mathbb{P}(\sup_t E_t \ge 1/δ_e) \le δ_e + δ_{\mathrm{cal}}$,相同总预算下的霍夫丁拼接累积误覆盖包络,对任何可预测控制器(重校准、带宽提升、学生刷新)的安全性,以及通过可累加训练预算实现无限次联邦探针逻辑蒸馏(FPLD)刷新中的误差传播控制。实际应用中,仅在e-过程越过预警阈值时才提升检索带宽,即可达到固定高带宽方案的报警率,但通信成本显著降低。在GPT-2-small + MiniLM集群上对MMLU、DBpedia和AG News的实验验证了预测的报警率、检测延迟、包络覆盖率及14%-57%的带宽节省;报警仅在覆盖真正失效时触发。

原文摘要 · Abstract (English)

Federated Conformal RAG (FC-RAG) provides distribution-free coverage for a bandwidth-limited swarm of weak language models, but only at a fixed horizon. We extend it to anytime-valid sequential coverage: validity at every stopping time, preserved under predictable adaptive control (recalibration, per-node bandwidth escalation, distilled-student refresh), at no extra cost in assumptions over fixed-horizon FC-RAG. Naive composition fails because FC-RAG's marginal coverage bound makes the betting e-process a non-supermartingale on adverse calibration draws, and Ville's inequality cannot be invoked. We give Anytime-FC-RAG, a sequential extension built on a summable per-step calibration-deviation budget that converts the marginal bound into a strict conditional bound on a calibration-good event, paired with a truncated betting e-process that is a nonnegative supermartingale on the entire probability space. From these two ingredients, we obtain four guarantees: time-uniform alarm validity $\mathbb{P}(\sup_t E_t \ge 1/δ_e) \le δ_e + δ_{\mathrm{cal}}$, a Hoeffding-stitched cumulative-miscoverage envelope at the same total budget, safety under any predictable controller (recalibration, bandwidth escalation, student refresh), and training-side error propagation across an unbounded sequence of Federated Probe-Logit Distillation (FPLD) refreshes via a summable training budget. As a practical consequence, an adaptive controller that escalates retrieval bandwidth only when the e-process crosses a warning threshold matches the alarm rate of a fixed-high-bandwidth schedule at substantially lower communication cost. Experiments on a GPT-2-small + MiniLM swarm across MMLU, DBpedia, and AG News verify the predicted alarm rate, detection delay, envelope coverage, and $14$-$57\%$ bandwidth savings; the alarm fires when and only when coverage genuinely breaks.

联邦学习可靠性高效通信语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。