arXiv:2412.04060cs.AI2024-12

通过选择性融合多源知识,提升异构系统下个性化模型的性能与效率。

Expand Heterogeneous Learning Systems with Selective Multi-Source Knowledge Fusion

  • 低成本筛选高质量模型,按样本权重融合其预测结果。
  • 在多任务多模态实验中,最高提升16.5%准确率,通信量减少39%。
  • 适合需要跨域个性化建模且资源受限的场景,如边缘计算。

将现有学习系统扩展至新领域(如新用户)以提供高质量定制化模型,面临标注数据有限及数据与设备异构性的挑战。尽管知识蒸馏可缓解标签稀缺与设备异构问题,但其假设教师模型完全可靠,忽视了数据异构性,导致无法直接应用现有模型。为此,本文提出框架HaT:首先低成本筛选系统中多个高质量模型,再通过分配样本级权重融合其预测;随后根据知识质量选择性地将融合知识注入定制模型。在不同任务、模态和设置下的大量实验表明,HaT相比最先进基线最高提升16.5%准确率,通信流量最多节省39%。

原文摘要 · Abstract (English)

Expanding existing learning systems to provide high-quality customized models for more domains, such as new users, is challenged by the limited labeled data and the data and device heterogeneities. While knowledge distillation methods could overcome label scarcity and device heterogeneity, they assume the teachers are fully reliable and overlook the data heterogeneity, which prevents the direct adoption of existing models. To address this problem, this paper proposes a framework, HaT, to expand learning systems. It first selects multiple high-quality models from the system at a low cost and then fuses their knowledge by assigning sample-wise weights to their predictions. Later, the fused knowledge is selectively injected into the customized models based on the knowledge quality. Extensive experiments on different tasks, modalities, and settings show that HaT outperforms state-of-the-art baselines by up to 16.5% accuracy and saves up to 39% communication traffic.

知识蒸馏异构系统个性化建模联邦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。