arXiv:2503.04183cs.LGcs.AI2025-03中稿 · IEEE Transactions …被引 3

动态自适应中间件,让手机端深度学习模型自动适配环境变化。

CrowdHMTware: A Cross-level Co-adaptation Middleware for Context-aware Mobile DL Deployment

  • 构建跨层级协同的自动化反馈机制,实现推理、卸载与模型的联动优化。
  • 在15种设备上测试,支持多种任务,显著提升模型部署的适应性与稳定性。
  • 适合需要低门槛部署AI应用的开发者,尤其在复杂多变的移动环境中。

当前众多移动与可穿戴应用依赖深度学习(DL)持续感知环境以改善人类生活。为保障移动端传感的鲁棒性与隐私性,常将DL模型部署于资源受限的本地设备,采用模型压缩或卸载技术。然而,现有方法要么在前端算法层面(如模型压缩/分割),要么在后端调度层面(如操作/资源调度),难以实现本地在线自适应:前者需离线重训练以保证精度,后者依赖人工预设策略,缺乏动态调整能力。核心挑战在于如何将后端运行时性能反馈至前端优化决策。此外,具备跨层级协同适应能力的移动DL模型部署中间件仍较少被研究,尤其是在多样且动态的移动环境中。为此,我们提出CrowdHMTware,一种面向异构移动设备的动态上下文自适应深度学习模型部署中间件。它建立了弹性推理、可扩展卸载与模型自适应引擎之间的自动化适应闭环,提升了系统的可扩展性与适应性。在四个典型任务及15个平台上的实验,以及一个真实场景案例研究均表明,CrowdHMTware能有效协调模型、卸载与引擎动作,在多样平台与任务中实现高效部署。它屏蔽了运行时系统问题,降低开发者所需专业知识门槛。

原文摘要 · Abstract (English)

There are many deep learning (DL) powered mobile and wearable applications today continuously and unobtrusively sensing the ambient surroundings to enhance all aspects of human lives.To enable robust and private mobile sensing, DL models are often deployed locally on resource-constrained mobile devices using techniques such as model compression or offloading.However, existing methods, either front-end algorithm level (i.e. DL model compression/partitioning) or back-end scheduling level (i.e. operator/resource scheduling), cannot be locally online because they require offline retraining to ensure accuracy or rely on manually pre-defined strategies, struggle with dynamic adaptability.The primary challenge lies in feeding back runtime performance from the back-end level to the front-end level optimization decision. Moreover, the adaptive mobile DL model porting middleware with cross-level co-adaptation is less explored, particularly in mobile environments with diversity and dynamics. In response, we introduce CrowdHMTware, a dynamic context-adaptive DL model deployment middleware for heterogeneous mobile devices. It establishes an automated adaptation loop between cross-level functional components, i.e. elastic inference, scalable offloading, and model-adaptive engine, enhancing scalability and adaptability. Experiments with four typical tasks across 15 platforms and a real-world case study demonstrate that CrowdHMTware can effectively scale DL model, offloading, and engine actions across diverse platforms and tasks. It hides run-time system issues from developers, reducing the required developer expertise.

移动AI自适应部署中间件边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。