arXiv:2509.25195cs.SEcs.LG2025-09被引 2

91名从业者调研揭示机器学习监控痛点与改进方向

Understanding Practitioners Perspectives on Monitoring Machine Learning Systems

  • 通过全球调研分析实践者在模型监控中的真实做法
  • 性能下降、延迟超标和安全违规是主要运行时问题
  • 亟需自动化监控生成与故障诊断建议功能

由于机器学习系统固有的非确定性,其在生产环境中的行为可能导致意外且潜在危险的结果。为及时发现异常行为并防止组织遭受财务和声誉损失,监控至关重要。本文从实践者视角探索了机器学习系统监控的策略、挑战与改进机会。我们对91名机器学习从业者进行了全球调查,收集了当前监控实践的多样见解。研究聚焦于常见的运行时问题、工业界的监控与缓解措施、关键挑战以及未来监控工具的期望改进。结果显示,实践者常面临模型性能下降、延迟超标和安全违规等运行时问题。尽管多数人偏好自动化监控以提高效率,但仍因复杂性或缺乏合适自动化解决方案而依赖手动方式。监控工具的初始配置和集成困难,特别是设置告警阈值时尤为棘手。此外,监控增加了额外工作量,耗用资源并导致告警疲劳。实践者期望的改进包括:监控的自动创建与部署、性能与公平性监控支持增强,以及运行时问题的修复建议。这些洞察为开发更契合实际需求的机器学习监控工具提供了重要指导。

原文摘要 · Abstract (English)

Given the inherent non-deterministic nature of machine learning (ML) systems, their behavior in production environments can lead to unforeseen and potentially dangerous outcomes. For a timely detection of unwanted behavior and to prevent organizations from financial and reputational damage, monitoring these systems is essential. This paper explores the strategies, challenges, and improvement opportunities for monitoring ML systems from the practitioners perspective. We conducted a global survey of 91 ML practitioners to collect diverse insights into current monitoring practices for ML systems. We aim to complement existing research through our qualitative and quantitative analyses, focusing on prevalent runtime issues, industrial monitoring and mitigation practices, key challenges, and desired enhancements in future monitoring tools. Our findings reveal that practitioners frequently struggle with runtime issues related to declining model performance, exceeding latency, and security violations. While most prefer automated monitoring for its increased efficiency, many still rely on manual approaches due to the complexity or lack of appropriate automation solutions. Practitioners report that the initial setup and configuration of monitoring tools is often complicated and challenging, particularly when integrating with ML systems and setting alert thresholds. Moreover, practitioners find that monitoring adds extra workload, strains resources, and causes alert fatigue. The desired improvements from the practitioners perspective are: automated generation and deployment of monitors, improved support for performance and fairness monitoring, and recommendations for resolving runtime issues. These insights offer valuable guidance for the future development of ML monitoring tools that are better aligned with practitioners needs.

机器学习监控实践调研运维挑战

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。