将智能推理嵌入云平台,实现自适应运维。
Cognitive Platform Engineering for Autonomous Cloud Operations
- 构建四层架构,融合感知、推理与自动执行
- 缩短故障修复时间,提升资源效率与合规性
- 适合追求自动化运维的云平台团队
现代DevOps通过自动化、CI/CD流水线和可观测性工具加速了软件交付,但面对云原生系统的规模与动态性仍显不足。随着遥测数据量增长和配置漂移加剧,传统规则驱动的自动化常导致被动响应、修复延迟及对人工经验的依赖。本文提出认知平台工程,一种将感知、推理与自主行动直接融入平台生命周期的新范式。提出四平面参考架构,整合数据采集、智能推断、策略编排与人机体验层,并形成持续反馈闭环。基于Kubernetes、Terraform、Open Policy Agent与基于机器学习的异常检测构建原型,验证了在平均故障修复时间、资源效率与合规性方面的提升。结果表明,将智能嵌入平台运维可实现弹性、自适应且意图对齐的云环境。最后展望强化学习、可解释治理与可持续自管理云生态等研究方向。
原文摘要 · Abstract (English)
Modern DevOps practices have accelerated software delivery through automation, CI/CD pipelines, and observability tooling,but these approaches struggle to keep pace with the scale and dynamism of cloud-native systems. As telemetry volume grows and configuration drift increases, traditional, rule-driven automation often results in reactive operations, delayed remediation, and dependency on manual expertise. This paper introduces Cognitive Platform Engineering, a next-generation paradigm that integrates sensing, reasoning, and autonomous action directly into the platform lifecycle. This paper propose a four-plane reference architecture that unifies data collection, intelligent inference, policy-driven orchestration, and human experience layers within a continuous feedback loop. A prototype implementation built with Kubernetes, Terraform, Open Policy Agent, and ML-based anomaly detection demonstrates improvements in mean time to resolution, resource efficiency, and compliance. The results show that embedding intelligence into platform operations enables resilient, self-adjusting, and intent-aligned cloud environments. The paper concludes with research opportunities in reinforcement learning, explainable governance, and sustainable self-managing cloud ecosystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。