跨数据集评估汽车网络入侵检测系统,发现性能差异大,强调测试需更全面。
CAN We Trust Your Results? A Cross-Dataset Study of Automotive IDS Evaluation

- 构建统一框架,整合7个不同条件下的汽车CAN网络数据集
- 5种不同方法在跨数据集测试中表现波动显著,最高差距超30%
- 适合研究车载安全与评估方法可靠性的开发者参考
现代车辆连接性增强,保障车内通信网络成为关键挑战。入侵检测系统(IDS)被广泛研究用于检测控制器局域网(CAN)总线上的恶意行为。然而,由于实验设置不一致且缺乏标准化评估框架,现有CAN IDS方法的性能评估存在困难,导致结果易受特定数据集特征影响,难以反映真实环境中的泛化能力。本文提出一个跨数据集的基准评估框架,整合了7个在不同实验条件下收集的公开CAN IDS数据集,对5种具有不同原理的IDS方法进行跨数据集测试。结果表明,检测性能在不同数据集间存在显著差异,凸显了跨数据集评估在衡量系统鲁棒性与泛化能力方面的重要性。
原文摘要 · Abstract (English)
The increasing connectivity of modern vehicles has made securing in-vehicle communication networks a critical challenge. Intrusion Detection Systems (IDS) have been widely studied as a defense mechanism for detecting malicious activities on the Controller Area Network (CAN) bus. However, the evaluation of CAN IDS methods remains difficult due to inconsistencies in experimental setups and the lack of standardized benchmarking frameworks. As a result, reported performance often depends on dataset-specific characteristics and may not reflect how detection methods behave in different environments. This work introduces a benchmarking framework for consistent evaluation of CAN IDSs across multiple datasets. Using the proposed framework, we integrate seven publicly available CAN IDS datasets collected under different experimental conditions and perform cross-dataset evaluation of five conceptually different IDS approaches. Our results highlight how detection performance can vary significantly across datasets, demonstrating the importance of cross-dataset benchmarking for assessing the robustness and generalization capabilities of CAN IDS methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。