用教师指导提升网络自主操作的强化学习训练效率。
A Comparative Evaluation of Teacher-Guided Reinforcement Learning Techniques for Autonomous Cyber Operations
- 引入四种教师引导技术,加速智能体在网络安全环境中的学习。
- 教师引导使早期策略表现和收敛速度显著提升。
- 适合关注自动化网络安全决策的研宄者与工程师。
自主网络操作(ACO)依赖强化学习(RL)训练智能体在网络安全领域做出有效决策。然而,现有ACO应用要求智能体从零开始学习,导致收敛缓慢且初期性能差。尽管教师引导技术在其他领域已证明有效,尚未应用于ACO。本研究在模拟的CybORG环境中实现并比较了四种不同的教师引导技术。结果表明,引入教师可显著提升训练效率,体现在早期策略性能和收敛速度上,凸显其在自主网络安全中的潜在优势。
原文摘要 · Abstract (English)
Autonomous Cyber Operations (ACO) rely on Reinforcement Learning (RL) to train agents to make effective decisions in the cybersecurity domain. However, existing ACO applications require agents to learn from scratch, leading to slow convergence and poor early-stage performance. While teacher-guided techniques have demonstrated promise in other domains, they have not yet been applied to ACO. In this study, we implement four distinct teacher-guided techniques in the simulated CybORG environment and conduct a comparative evaluation. Our results demonstrate that teacher integration can significantly improve training efficiency in terms of early policy performance and convergence speed, highlighting its potential benefits for autonomous cybersecurity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。