arXiv:2512.23043cs.LG2025-12

非独立同分布下联邦学习模型仍保留有用特征,但特征与预测路径错位导致性能下降。

Mechanistic Evidence for Preserved-but-Misaligned Representations in Non-IID FedAvg

  • 通过电路发现与线性探测,验证稀疏模型中类特定结构仍可恢复
  • 在严重标签偏斜下,部分类别准确率接近零但内部结构未丢失
  • 适合关注联邦学习表征对齐问题的研究者阅读

联邦平均(FedAvg)在客户端数据非独立同分布(non-IID)时性能常下降,但尚不清楚这是由于客户端学到的表征丢失,还是现有表征未能被有效利用。本文在稀疏训练的视觉模型中进行机制分析,使用密集模型作为对照,检验稀疏性是否影响结果。基于类特定电路发现、冻结表征的线性探测、仅微调头部和稀疏特征字典等方法,在CIFAR-10和Fashion-MNIST上的CNN与ResNet模型中发现:严重标签偏斜会导致某些类别的准确率趋近于零,但类特定内部结构仍可恢复。线性探测性能显著优于聚合分类器,仅微调头部可部分恢复准确率,且USAE迁移分析显示,IID与non-IID模型间存在广泛共享的特征基础。这些诊断表明,在本设置中,non-IID FedAvg性能下降并非完全由表征消退引起,更关键的是保留的内部结构与最终预测路径之间的错位。

原文摘要 · Abstract (English)

Federated Averaging (FedAvg) often degrades under non-IID client data, but it remains unclear whether this degradation reflects the loss of client-learned representations or a failure to use representations that are still present. We study this question mechanistically in sparse client-trained vision models, using dense-model controls to test whether the observed effects depend on sparsity. Our analysis combines class-specific circuit discovery, linear probing of frozen representations, head-only finetuning, and sparse feature dictionaries. Across CNN and ResNet models on CIFAR-10 and Fashion-MNIST, severe label skew can drive some per-class accuracies near zero even when class-specific internal structure remains recoverable. Linear probes substantially outperform the aggregated classifier, head-only finetuning partially restores accuracy, and USAE transfer reveals a largely shared feature basis between IID and non-IID models. Together, these diagnostics suggest that, in our setting, non-IID FedAvg degradation is not fully explained by representational erasure; it also reflects misalignment between preserved internal structure and the final prediction pathway.

联邦学习表征对齐非IID稀疏模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。