arXiv:2509.05311cs.CRcs.AI2025-09被引 6

用大模型指导强化学习,让网络自主作战更安全高效。

Large Language Model Integration with Reinforcement Learning to Augment Decision-Making in Autonomous Cyber Operations

  • 用预训练的网络安全大模型引导强化学习初始决策
  • 早期训练奖励提升2倍,收敛速度加快4500回合
  • 适合需要快速可靠决策的网络安全自动化场景

强化学习在自主网络作战中展现出巨大潜力,使智能体通过与环境交互自主学习决策。然而,当前自主网络作战中的强化学习智能体通常从零开始学习,需执行不当操作以了解后果,存在安全风险。本研究将预训练于网络安全数据的大语言模型(LLM)作为外部知识源,使强化学习智能体可直接利用该知识进行决策。通过用大模型指导初始训练,显著提升了基线性能,并减少了具有明显负面后果的探索性行为。我们在模拟网络安全环境中评估了该方法,结果表明,受大模型引导的智能体在早期训练中获得超过2倍的奖励,且比基线智能体提前约4500个训练回合收敛至优良策略。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has shown great potential for autonomous decision-making in the cybersecurity domain, enabling agents to learn through direct environment interaction. However, RL agents in Autonomous Cyber Operations (ACO) typically learn from scratch, requiring them to execute undesirable actions to learn their consequences. In this study, we integrate external knowledge in the form of a Large Language Model (LLM) pretrained on cybersecurity data that our RL agent can directly leverage to make informed decisions. By guiding initial training with an LLM, we improve baseline performance and reduce the need for exploratory actions with obviously negative outcomes. We evaluate our LLM-integrated approach in a simulated cybersecurity environment, and demonstrate that our guided agent achieves over 2x higher rewards during early training and converges to a favorable policy approximately 4,500 episodes faster than the baseline.

强化学习大模型网络安全自主决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。