用JEPA方法学习网络指纹,效果接近92%准确率。
Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints
- 用Transformer模型基于JA4数据做预测性自监督学习
- 在3.9万样本上达0.922的分类准确率和0.99相似度
- 适合研究网络指纹与自监督学习的学者
I-JEPA和V-JEPA通过匹配潜在预测与目标编码器输出来学习,而非重建原始输入,在图像和视频中表现良好。本文探索该方法是否适用于紧凑的网络指纹。构建了基于Transformer的JA4-JEPA模型,训练数据来自JA4DB和CIC-IDS-2017中的JA4、JA4H、JA4S、JA4X子集,共约39.7万个样本,但无单一样本包含全部四类视图。在冻结kNN探测器下,对TLS、DNS、SSH协议族分类任务进行评估,使用39,416个保留样本,模型达到0.9899的余弦相似度和0.9220的kNN准确率。结果表明,即使跨源视图重叠不全,JEPA式预测学习仍可生成有效的JA4衍生指纹嵌入。
原文摘要 · Abstract (English)
I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has worked well for images and video. We explore whether the same objective works for compact network fingerprints. We built JA4-JEPA, a Transformer-based model trained on JA4, JA4H, JA4S, and JA4X subfields drawn from JA4DB and CIC-IDS- 2017. The training data combines roughly 397K samples from both sources, though no single sample contains all four view families. We evaluated the learned representations with a frozen kNN probe on protocol-family classification across TLS, DNS, and SSH. On 39,416 heldout samples the model achieved a cosine similarity of 0.9899 and a kNN accuracy of 0.9220. These results indicate that JEPA-style predictive learning can produce useful embeddings from JA4-derived fingerprints, even with incomplete view overlap across sources. Keywords: JA4, network fingerprinting, JEPA, predictive representation learning, self-supervised learning
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。