arXiv:2510.26008cs.PFcs.AR2025-10

通过硬件信号检测机器学习系统异常,无需了解具体模型即可优化性能。

Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry

  • 仅依赖硬件信号,用无监督学习检测系统异常
  • 在多种硬件上验证,成功提升DeepSeek模型5.97%速度
  • 适合云平台运维人员快速发现配置问题

现代机器学习已发展为软硬件紧密耦合的全栈生态系统,用户依赖云服务商提供弹性、隔离且成本可控的资源。然而,这些平台即服务(PaaS)普遍采用虚拟化技术,导致运营商对用户工作负载知之甚少,难以进行资源优化,影响成本效率与执行时长。本文提出Reveal,一种以硬件为中心的系统级异常检测方法,仅依赖运营商可访问的硬件信号。基于在多种硬件平台上对30余种主流ML模型的分析,构建了无监督学习管道,具备对新兴工作负载和未知部署模式的适应能力。实验表明,Reveal成功识别出网络与系统配置问题,使DeepSeek模型推理速度提升5.97%。

原文摘要 · Abstract (English)

Modern machine learning (ML) has grown into a tightly coupled, full-stack ecosystem that combines hardware, software, network, and applications. Many users rely on cloud providers for elastic, isolated, and cost-efficient resources. Unfortunately, these platforms as a service use virtualization, which means operators have little insight into the users' workloads. This hinders resource optimizations by the operator, which is essential to ensure cost efficiency and minimize execution time. In this paper, we argue that workload knowledge is unnecessary for system-level optimization. We propose Reveal, which takes a hardware-centric approach, relying only on hardware signals - fully accessible by operators. Using low-level signals collected from the system, Reveal detects anomalies through an unsupervised learning pipeline. The pipeline is developed by analyzing over 30 popular ML models on various hardware platforms, ensuring adaptability to emerging workloads and unknown deployment patterns. Using Reveal, we successfully identified both network and system configuration issues, accelerating the DeepSeek model by 5.97%.

系统监控异常检测硬件信号云平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。