让联邦学习学会动态优化规则,实现零样本测试时自适应。
Federated Nested Learning: Collaborative Training of Self-Referential Memories for Test-Time Adaptation

- 将联邦学习重构为三层嵌套优化,协作学习动态优化策略。
- 在非独立同分布数据上,短文本推理性能达基准水平,长文本检索显著提升。
- 轻量级设计支持零样本测试时调整,适合资源受限设备部署。
我们从嵌套学习视角重新思考联邦学习(FL),将核心挑战定位为如何协同学习优化规则而非静态模型,以应对客户端数据的非独立同分布问题。为此,提出联邦嵌套学习(FedNL)框架,将联邦学习重构为三层嵌套优化系统。FedNL引入基于Titans的线性注意力机制,使客户端能够通过将增量规则视为在线梯度步,实现轻量级、零样本的测试时自适应。在非独立同分布的MMLU和长上下文基准上的实验表明,FedNL在短文本推理中达到竞争力表现,显著提升长文本检索与流式交叉熵任务性能,且推理内存保持恒定。
原文摘要 · Abstract (English)
We rethink Federated Learning (FL) from a nested learning perspective, framing the core challenge as how to collaboratively learn optimization rules, not just static models, to tackle Non-IID client data. To address this, we propose Federated Nested Learning (FedNL), a novel framework that reformulates FL as a three-level nested optimization system. FedNL embeds Titans-based linear attention into FL, enabling clients to perform lightweight, zero-shot test-time adaptation by treating a delta rule as an online gradient step. Experiments on Non-IID MMLU and long-context benchmarks show that FedNL achieves competitive performance in short-context reasoning, enhances the performance of long-context retrieval and streaming Cross-Entropy, and maintains constant inference memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。