解决车载攻击数据稀缺问题,生成高仿真攻击数据用于安全检测
Alleviating Attack Data Scarcity: SCANIA's Experience Towards Enhancing In-Vehicle Cyber Security Measures
- 基于参数化模型生成多种攻击的CAN网络日志
- 生成数据使深度神经网络检测模型达到高检出率
- 适合车联网安全研究与检测系统开发人员
联网汽车的数字化发展带来安全风险,亟需部署入侵检测与响应系统。随着攻击场景不断演进,需具备识别未知复杂威胁的自适应检测机制。尽管机器学习技术可应对挑战,但受限于安全、成本与伦理,真实测试车辆难以复现多样攻击场景,导致攻击数据稀缺。本文提出一种上下文感知的攻击数据生成器,可生成包括拒绝服务、模糊、伪造、暂停、重放等攻击类型的输入及对应的控制器局域网(CAN)日志。通过参数化攻击模型结合CAN消息解析与攻击强度调节,实现与真实场景高度相似且具有多样性的攻击配置。在入侵检测系统(IDS)案例研究中,使用生成数据训练并评估两个深度神经网络模型,结果表明其具备高效性、可扩展性以及优异的检测与分类能力,验证了生成数据的一致性与有效性。研究还分析了影响数据真实度的关键因素,并提供实际应用建议。
原文摘要 · Abstract (English)
The digital evolution of connected vehicles and the subsequent security risks emphasize the critical need for implementing in-vehicle cyber security measures such as intrusion detection and response systems. The continuous advancement of attack scenarios further highlights the need for adaptive detection mechanisms that can detect evolving, unknown, and complex threats. The effective use of ML-driven techniques can help address this challenge. However, constraints on implementing diverse attack scenarios on test vehicles due to safety, cost, and ethical considerations result in a scarcity of data representing attack scenarios. This limitation necessitates alternative efficient and effective methods for generating high-quality attack-representing data. This paper presents a context-aware attack data generator that generates attack inputs and corresponding in-vehicle network log, i.e., controller area network (CAN) log, representing various types of attack including denial of service (DoS), fuzzy, spoofing, suspension, and replay attacks. It utilizes parameterized attack models augmented with CAN message decoding and attack intensity adjustments to configure the attack scenarios with high similarity to real-world scenarios and promote variability. We evaluate the practicality of the generated attack-representing data within an intrusion detection system (IDS) case study, in which we develop and perform an empirical evaluation of two deep neural network IDS models using the generated data. In addition to the efficiency and scalability of the approach, the performance results of IDS models, high detection and classification capabilities, validate the consistency and effectiveness of the generated data as well. In this experience study, we also elaborate on the aspects influencing the fidelity of the data to real-world scenarios and provide insights into its application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。