用可验证奖励强化学习,让大模型更准地生成安全威胁情报。
Minerva: Reinforcement Learning with Verifiable Rewards for Cyber Threat Intelligence LLMs
- 通过可验证反馈机制,让模型输出能自动核对正确性。
- 在12个威胁情报任务上平均提升15.8个百分点性能。
- 适合安全分析、自动化系统开发人员使用。
网络安全分析师需将杂乱无章的安全数据转化为标准化、可自动处理的格式。尽管大语言模型(LLMs)在此任务中展现出潜力,但现有方法在生成结构化威胁情报时仍不稳定,且主要依赖监督微调(SFT)。然而,威胁情报标准与社区维护资源提供了可确定性验证的标识符和模式。我们利用这一特性,探索用于威胁情报任务的可验证奖励强化学习(RLVR)。提出Minerva:一个统一的数据集与训练流程,涵盖多个子任务,并配备任务专用验证器,用于评估结构化输出与标识符预测。为缓解采样过程中的奖励稀疏问题,提出MinervaRL——一种轻量级自训练机制,可生成额外已验证轨迹并回传至模型。在四种骨干模型与12个威胁情报基准上的平均结果显示,MinervaRL相较基础模型提升15.8个百分点,较GRPO提升4.3个百分点。
原文摘要 · Abstract (English)
Cyber threat intelligence (CTI) analysts routinely convert noisy, unstructured security artifacts into standardized, automation-ready representations. Although large language models (LLMs) show promise for this task, existing approaches remain brittle when producing structured CTI outputs and have largely relied on supervised fine-tuning (SFT). In contrast, CTI standards and community-maintained resources define canonical identifiers and schemas that enable deterministic verification of model outputs. We leverage this structure to study reinforcement learning with verifiable rewards (RLVR) for CTI tasks. We introduce Minerva, a unified dataset and training pipeline spanning multiple CTI subtasks, each paired with task-specific verifiers that score structured outputs and identifier predictions. To address reward sparsity during rollout, we propose MinervaRL, a lightweight self-training mechanism that generates additional verified trajectories and distills them back into the model. Averaged across four backbones and 12 CTI benchmarks, MinervaRL improves the mean score by 15.8 percentage points over the corresponding base models and by 4.3 points over GRPO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。