构建统一基准,系统评估生成式密码猜测模型性能。
MAYA: Addressing Inconsistencies in Generative Password Guessing through a Unified Benchmark
- 提出可插拔的统一评估框架MAYA,支持标准化测试。
- 在8个真实数据集上测试6种模型,耗时超15000小时,发现序列模型更优。
- 揭示长复杂密码下性能下降,多模型协同攻击效果最佳。
生成式模型在密码猜测中的应用日益广泛,旨在复现人类创建密码的复杂性与规律。然而,现有研究存在评估方法不一致、结果不可比等问题,阻碍了对模型能力的客观理解。本文提出MAYA——一个统一、可定制、即插即用的基准框架,用于系统化评估生成式密码猜测模型在拖网攻击场景下的表现。通过MAYA,我们对六种前沿方法进行了重新实现与标准化适配,覆盖8个真实世界密码数据集,涵盖超过15,000小时的计算量。结果表明,这些模型能有效捕捉人类密码分布的不同特征,具备较强泛化能力,但在长而复杂的密码上表现差异显著。序列模型始终优于其他生成架构及传统工具,展现出生成精准复杂猜测的独特能力。此外,各模型学习到的多样化密码分布,使多模型联合攻击优于单一最优模型。MAYA已开源,旨在推动社区研究,提供可靠、一致的评估工具。
原文摘要 · Abstract (English)
Recent advances in generative models have led to their application in password guessing, with the aim of replicating the complexity, structure, and patterns of human-created passwords. Despite their potential, inconsistencies and inadequate evaluation methodologies in prior research have hindered meaningful comparisons and a comprehensive, unbiased understanding of their capabilities. This paper introduces MAYA, a unified, customizable, plug-and-play benchmarking framework designed to facilitate the systematic characterization and benchmarking of generative password-guessing models in the context of trawling attacks. Using MAYA, we conduct a comprehensive assessment of six state-of-the-art approaches, which we re-implemented and adapted to ensure standardization. Our evaluation spans eight real-world password datasets and covers an exhaustive set of advanced testing scenarios, totaling over 15,000 compute hours. Our findings indicate that these models effectively capture different aspects of human password distribution and exhibit strong generalization capabilities. However, their effectiveness varies significantly with long and complex passwords. Through our evaluation, sequential models consistently outperform other generative architectures and traditional password-guessing tools, demonstrating unique capabilities in generating accurate and complex guesses. Moreover, the diverse password distributions learned by the models enable a multi-model attack that outperforms the best individual model. By releasing MAYA, we aim to foster further research, providing the community with a new tool to consistently and reliably benchmark generative password-guessing models. Our framework is publicly available at https://github.com/williamcorrias/MAYA-Password-Benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。