arXiv:2410.18378cs.LG2024-10被引 12

用云端数据补足设备端数据短板,提升手机持续学习效果

Delta: A Cloud-assisted Data Enrichment Framework for On-Device Continual Learning

  • 分阶段处理:设备端与云端协作,不共享敏感数据
  • 平均提升准确率15.1%~12.4%,通信开销降低90%以上
  • 适合资源受限的移动端持续学习场景

在现代移动应用中,用户频繁遭遇新场景,需依赖设备端持续学习(CL)以维持模型性能。现有研究多聚焦轻量化框架,我们发现数据稀缺是设备端持续学习的关键瓶颈。本文探索利用云端丰富数据补充设备端稀缺数据,提出隐私、高效、有效的数据增强框架Delta。首先引入目录数据集,将数据增强问题分解为设备端与云端子任务,避免敏感数据泄露;其次提出软数据匹配策略,有效解决设备端稀疏用户数据的问题;设计低复杂度最优数据采样方案,从云端检索最适配的数据集进行增强;进一步联合考虑新增与旧有场景的影响,优化采样策略,缓解灾难性遗忘。在四种典型移动端任务(视觉、IMU、音频、文本)上实验表明,相比少样本持续学习,Delta平均提升准确率15.1%、12.4%、1.1%和5.6%;相比联邦学习,通信开销始终降低90%以上。

原文摘要 · Abstract (English)

In modern mobile applications, users frequently encounter various new contexts, necessitating on-device continual learning (CL) to ensure consistent model performance. While existing research predominantly focused on developing lightweight CL frameworks, we identify that data scarcity is a critical bottleneck for on-device CL. In this work, we explore the potential of leveraging abundant cloud-side data to enrich scarce on-device data, and propose a private, efficient and effective data enrichment framework Delta. Specifically, Delta first introduces a directory dataset to decompose the data enrichment problem into device-side and cloud-side sub-problems without sharing sensitive data. Next, Delta proposes a soft data matching strategy to effectively solve the device-side sub-problem with sparse user data, and an optimal data sampling scheme for cloud server to retrieve the most suitable dataset for enrichment with low computational complexity. Further, Delta refines the data sampling scheme by jointly considering the impact of enriched data on both new and past contexts, mitigating the catastrophic forgetting issue from a new aspect. Comprehensive experiments across four typical mobile computing tasks with varied data modalities demonstrate that Delta could enhance the overall model accuracy by an average of 15.1%, 12.4%, 1.1% and 5.6% for visual, IMU, audio and textual tasks compared with few-shot CL, and consistently reduce the communication costs by over 90% compared to federated CL.

持续学习云端协同数据增强移动计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。