将数据质量评估与模型运维结合,实现实时高效工业预测。
End-to-End Data Quality-Driven Framework for Machine Learning in Production Environment
- 动态检测数据漂移,自适应调整质量指标
- 提升模型性能至R2=94%,预测延迟降低75%
- 适合对实时性要求高的工业场景
本文提出一种端到端框架,将数据质量评估与机器学习系统运维在实时生产环境中无缝集成。现有方法常将两者割裂,本框架通过动态漂移检测、自适应质量指标与MLOps融合,构建轻量级一体化系统。核心优势在于操作高效,实现低开销的实时质量驱动决策。在某钢铁企业电渣重熔(ESR)真空泵送过程验证中,模型性能提升12%(R2=94%),预测延迟降低四倍。通过分析数据质量阈值影响,提供工业应用中质量标准与预测性能平衡的实用建议。该框架显著推进MLOps发展,为动态工业环境中的时敏数据决策提供可靠方案。
原文摘要 · Abstract (English)
This paper introduces a novel end-to-end framework that efficiently integrates data quality assessment with machine learning (ML) model operations in real-time production environments. While existing approaches treat data quality assessment and ML systems as isolated processes, our framework addresses the critical gap between theoretical methods and practical implementation by combining dynamic drift detection, adaptive data quality metrics, and MLOps into a cohesive, lightweight system. The key innovation lies in its operational efficiency, enabling real-time, quality-driven ML decision-making with minimal computational overhead. We validate the framework in a steel manufacturing company's Electroslag Remelting (ESR) vacuum pumping process, demonstrating a 12% improvement in model performance (R2 = 94%) and a fourfold reduction in prediction latency. By exploring the impact of data quality acceptability thresholds, we provide actionable insights into balancing data quality standards and predictive performance in industrial applications. This framework represents a significant advancement in MLOps, offering a robust solution for time-sensitive, data-driven decision-making in dynamic industrial environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。