量子机器学习在电力系统攻击检测中表现好坏,取决于评测设计而非模型本身。
Benchmarking Quantum Machine Learning for Power-System Attack Detection: Evaluation Choices Decide the Outcome Before the Models Do

- 评测方法选择直接影响结论,而非模型性能
- 不同数据划分方式使结果从0.905到0.594差异巨大
- 适合关注评测可靠性与实验设计的科研人员
针对电力系统网络攻击的机器学习检测器本身也是攻击目标,量子机器学习被提出用于此类场景。我们在公开的电力系统攻击数据集(Mississippi State/ORNL)上,对保真度核SVM和变分分类器与六种调优后的经典模型进行对比,评估其在白盒、迁移、基于决策的黑盒及投毒攻击下的表现。核心发现为方法论层面:评测结果由评估者的选择决定,而非模型能力。八个关键选择——六项评估协议、两项基准调优设置——任一调整均能改变结论。最大影响来自数据划分方式:行级协议得宏平均F1为0.905,而完整源文件保留时仅0.594;在维度受限条件下,量子模型与经典模型差距仅0.024,处于噪声水平。保真度核看似最稳健,但在直接攻击下性能从0.886降至0.064;误拟合代理模型制造出10倍不对称性;未种子化的黑盒攻击在多次重启间导致75%结果漂移。正向控制揭示准确率无差异根源在于标签本身,而非流程。我们提供每项选择的对照方案并开源种子化评测基准。
原文摘要 · Abstract (English)
Machine-learning detectors for power-system cyberattacks are themselves attack surfaces, and quantum machine learning has been proposed for them. We benchmark fidelity-kernel SVMs and variational classifiers against six tuned classical models on public power-system attack data (Mississippi State/ORNL), across white-box, transfer, decision-based black-box, and poisoning attacks. Our headline finding is methodological: the benchmark's answers are set by the evaluator's choices before the models. Eight choices -- six in the evaluation protocol, two in the tuning the benchmark itself runs -- each reversed or moved a conclusion at fixed models. The largest is the split: the row-level protocol scores 0.905 macro-F1 where holding whole source files out leaves 0.594, and in the capped matched-dimensionality regime the quantum arm sits within noise of chance with the classical arm 0.024 above it. A fidelity kernel looks most robust until attacked directly (retention 0.886 to 0.064); a mis-fitted surrogate manufactures a 10x asymmetry; an unseeded black-box attack moves 75% between restarts. A positive control explains the accuracy null: the labels, not the pipeline. We give the control that catches each choice and release the seeded benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。