用开源框架训练模型检测多智能体系统中的时间攻击模式。
Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models
- 基于OpenTelemetry追踪数据,用迭代QLoRA微调语言模型。
- 准确率从42.86%提升至74.29%,31.4点显著增长。
- 数据集与代码全开源,适合安全研发人员定制防御模型。
我们提出一种公开文档化的方法,通过OpenTelemetry追踪分析,微调语言模型以检测多智能体AI工作流中的时间攻击模式。数据集包含来自18个公开网络安全源的80,851个样本和35,026条合成OpenTelemetry追踪。在资源受限的ARM64硬件(NVIDIA DGX Spark)上,采用三次迭代的QLoRA微调并结合策略性增强。自定义基准测试准确率从42.86%提升至74.29%,统计显著提高31.4个百分点。针对性解决知识盲区的样本表现优于盲目扩展。主要贡献包括:(1) 多智能体协同攻击与合规违规的合成追踪生成方法;(2) 实证表明训练数据构成从根本上决定模型行为;(3) 在HuggingFace上完整开源数据集、训练脚本与评估基准。尽管实际部署需人工审核以应对误报,该工作首次建立可复现的框架,使从业者能根据自身威胁环境构建定制化智能体安全模型。
原文摘要 · Abstract (English)
We present an openly documented methodology for fine-tuning language models to detect temporal attack patterns in multi-agent AI workflows using OpenTelemetry trace analysis. We curate a dataset of 80,851 examples from 18 public cybersecurity sources and 35,026 synthetic OpenTelemetry traces. We apply iterative QLoRA fine-tuning on resource-constrained ARM64 hardware (NVIDIA DGX Spark) through three training iterations with strategic augmentation. Our custom benchmark accuracy improves from 42.86% to 74.29%, a statistically significant 31.4-point gain. Targeted examples addressing specific knowledge gaps outperform indiscriminate scaling. Key contributions include: (1) synthetic trace generation methodology for multi-agent coordination attacks and regulatory violations, (2) empirical evidence that training data composition fundamentally determines behavior, and (3) complete open release of datasets, training scripts, and evaluation benchmarks on HuggingFace. While practical deployment requires human oversight due to false positive rates, this work establishes the first reproducible framework enabling practitioners to build custom agentic security models adapted to their threat landscapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。