通过主动因果学习在线构建IT系统因果模型,实现低干扰精准建模。
Online Identification of IT Systems through Active Causal Learning
- 基于高斯过程回归与滚动干预策略,迭代估计系统变量间因果关系。
- 实验表明模型识别准确率高,对系统运行干扰极小。
- 适合自动化运维、故障诊断等需要实时因果推理的场景。
识别IT系统的因果模型是系统工程与运维的核心任务,可用于预测控制行为效果、优化操作、故障诊断和入侵检测,是实现网络与系统管理自动化的关键。传统方法依赖领域专家设计维护,但面对现代IT系统日益复杂的动态性已难以为继。本文提出首个原理严谨的在线数据驱动因果模型识别方法——主动因果学习。该方法基于高斯过程回归,通过滚动干预策略收集系统观测数据,迭代估计系统变量间的因果函数。理论上证明该方法在贝叶斯意义下最优,且能生成有效干预。在测试平台上验证显示,该方法可实现高精度因果模型识别,同时对系统运行干扰极小。
原文摘要 · Abstract (English)
Identifying a causal model of an IT system is fundamental to many branches of systems engineering and operation. Such a model can be used to predict the effects of control actions, optimize operations, diagnose failures, detect intrusions, etc., which is central to achieving the longstanding goal of automating network and system management tasks. Traditionally, causal models have been designed and maintained by domain experts. This, however, proves increasingly challenging with the growing complexity and dynamism of modern IT systems. In this paper, we present the first principled method for online, data-driven identification of an IT system in the form of a causal model. The method, which we call active causal learning, estimates causal functions that capture the dependencies among system variables in an iterative fashion using Gaussian process regression based on system measurements, which are collected through a rollout-based intervention policy. We prove that this method is optimal in the Bayesian sense and that it produces effective interventions. Experimental validation on a testbed shows that our method enables accurate identification of a causal system model while inducing low interference with system operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。