arXiv:2607.08784cs.LGcs.AI2026-07

HERO为联邦持续学习提供可比性更强的异构性基准测试库。

HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning

论文配图:HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning
图 1 · 摘自论文原文
  • 分离任务划分、客户端数据分布和任务顺序,解耦评估变量。
  • 在CIFAR-100和TinyImageNet上验证方法,发现平均准确率会掩盖弱客户端表现。
  • 支持图像与图结构任务,适用于研究领域漂移对模型的影响。

联邦持续学习(FCL)评估分布式客户端在不断变化的数据流中持续学习并保留已有知识的能力。现有评估因同时改变数据集、任务划分、客户端数据分布、任务顺序、主干网络、记忆假设和报告规则而难以比较。我们提出 extbf{HERO},一个面向异构性的联邦持续学习基准库。HERO通过解耦三个常被耦合的要素——任务划分、客户端数据划分和客户端任务序列——构建可比的评估流。在核心基准HERO-Core中,参数 $α$ 控制客户端数据偏斜,$ρ$ 控制任务顺序不匹配。我们在CIFAR-100和TinyImageNet上使用最终平均准确率、平均遗忘率和底10%客户端准确率评估代表性FCL方法。还包含一个基于图的域独立学习可迁移性案例研究,针对OGB-MolPCBA,其中骨架域粒度改变输入分布但预测任务保持不变。结果表明:方法行为在简单与异构设置间差异显著;平均准确率可能掩盖弱客户端性能;任务顺序不匹配使策略选择不同于同步评估;同一HERO接口可揭示图像外的任务领域漂移难度。HERO发布基准流、配置、方法实现和报告脚本,支持可复现且设置感知的FCL评估。

原文摘要 · Abstract (English)

Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficult to compare because they often change datasets, task splits, client data splits, task orders, backbones, memory assumptions, and reporting rules simultaneously. We introduce \textbf{HERO}, a heterogeneity-aware benchmark library for FCL. HERO builds benchmark streams by separating three choices that are often coupled, namely the task split, the client data split, and the client task sequence. In HERO-Core, the main comparable benchmark, $α$ controls client data skew and $ρ$ controls task-order mismatch. We evaluate representative FCL methods on CIFAR-100 and TinyImageNet using final average accuracy, average forgetting, and bottom-10\% client accuracy. We also include a graph-based Domain-IL portability case study on OGB-MolPCBA, where scaffold-domain granularity changes the input distribution while the prediction task remains fixed. Our results show that method behavior changes across easy and heterogeneous settings, that average accuracy can hide weak bottom-client performance, that task-order mismatch favors different strategies from synchronized evaluation, and that the same HERO interface can expose domain-shift difficulty beyond image-based FCIL. HERO releases benchmark streams, configurations, method implementations, and reporting scripts to support reproducible and setting-aware FCL evaluation.

联邦学习持续学习基准测试异构性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。