不同测试划分方式显著影响硬件木马检测效果,兄弟变体共享逻辑会虚高性能。
Protocol effects on feature-based hardware-Trojan detection across Trust-Hub families
- 用同一类电路的多个变体作为训练和测试数据,模拟真实场景中的检测性能。
- 当测试时排除整个家族时,检测准确率从0.914降至0.460,差距巨大。
- 建议在报告中区分‘合并测试’与‘按家族隔离’结果,避免误导性结论。
Trust-Hub通过复用宿主电路生成多个变体,仅在其中插入木马。实验使用来自16个网表的49,124个门级单元,分为五个宿主家族。保持解析器、36个门特征、类别权重、模型设置、阈值和家族聚合方式不变,仅改变测试边界:从合并数据池中抽样、剔除一个完整网表,或剔除一个宿主家族的所有变体。随机森林在合并测试下F1/AP为0.914/0.978,剔除单个网表时降为0.636/0.851,剔除整个家族时更降至0.460/0.577;XGBoost也从0.946/0.976跌至0.464/0.544。逻辑回归虽保持固定阈值的F1不单调,但准确率(AP)下降。各家族均呈现相同趋势。特征移除、种子重复、分数归一化、解析排除及小样本等操作改变差距大小,但不反转方向。加权平均受大文件主导,因此采用每家族一票。自助法与刀切法验证仍维持正向差距,但其折叠重用训练家族。将五族视为描述性证据而非独立试验。因家族数不足,无法推广至新单元库或工业设计。结论是:兄弟变体可能虚增迁移性能。建议含多变体的基准报告应提供家族感知的留出策略及全部五族结果,而非仅汇报合并得分。
原文摘要 · Abstract (English)
Trust-Hub reuses host circuits: several files differ mainly in the inserted Trojan. When gates from sibling variants enter both training and test folds, a detector can benefit from host logic it has already seen. We measure that effect instead of proposing another classifier. The corpus contains 49,124 gates from 16 netlists grouped into five host families. We left the parser, 36 gate features, class weighting, model settings, threshold, and family-level aggregation unchanged and altered one choice: the test boundary. The three settings draw test gates from the pooled corpus, withhold a complete netlist, or withhold every variant of one host. The choice matters. Random forest records F1/AP of 0.914/0.978 with pooled gates, 0.636/0.851 with one netlist held out, and 0.460/0.577 with a host family held out. XGBoost falls from 0.946/0.976 to 0.464/0.544 across the same comparison. Logistic regression loses AP, although its fixed-threshold F1 is not monotonic. Each family shows the same pooled-to-family direction. Feature removal, repeated model and simulator seeds, score normalization, parser-related exclusions, and a smaller sample change the size of the gap without reversing it. Aggregation also matters: a gate-weighted average is dominated by the larger ISCAS files, so the headline values give each host family one vote. Bootstrap and jackknife summaries keep the gap positive, but their folds reuse training families. We treat the five family rows as descriptive evidence rather than independent trials. Five host families are too few for a population claim, and the experiment says nothing about transfer to a new cell library or an industrial design. It supports a narrower conclusion: sibling benchmark variants can inflate apparent transfer. Benchmarks with several variants of one host circuit should report family-aware holdouts and all five family results beside pooled scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。