用机器学习优化鼠标行为认证,解决数据少、准度与实用难兼顾问题。
Optimizing Mouse Dynamics for User Authentication by Machine Learning: Addressing Data Sufficiency, Accuracy-Practicality Trade-off, and Model Performance Challenges
- 用高斯核密度与KL散度估算训练所需最少数据量
- 通过近似熵确定最佳分段长度,提升识别效率与精度
- 新框架在数据量减10倍情况下仍超当前最佳性能
用户认证对保障系统安全至关重要,传统方法存在易用性差、成本高和安全性不足的问题。基于鼠标操作行为的动态认证提供了一种低成本、非侵入式且可适应的解决方案。然而,仍面临数据量不足、准确率与实用性难以平衡、时序行为模式捕捉困难等挑战。本研究提出一种基于高斯核密度估计(KDE)与Kullback-Leibler(KL)散度的统计方法,用于估算训练模型所需的足够数据量。引入鼠标认证单元(MAU),利用近似熵(ApEn)优化分段长度,实现高效准确的行为表征。设计局部时间鼠标认证(LT-AMouse)框架,结合1D-ResNet提取局部特征、GRU建模长期时序依赖。以Balabit和DFL数据集为例,显著降低数据规模,尤其使DFL数据集减少10倍,极大减轻训练负担。基于不同数据集的近似熵斜率,确定最优输入识别单元长度。在不平衡样本下训练,模型在盲攻击测试中对DFL数据集达到98.52%的AUC,Balabit数据集为94.65%,超越当前最先进水平。
原文摘要 · Abstract (English)
User authentication is essential to ensure secure access to computer systems, yet traditional methods face limitations in usability, cost, and security. Mouse dynamics authentication, based on the analysis of users' natural interaction behaviors with mouse devices, offers a cost-effective, non-intrusive, and adaptable solution. However, challenges remain in determining the optimal data volume, balancing accuracy and practicality, and effectively capturing temporal behavioral patterns. In this study, we propose a statistical method using Gaussian kernel density estimate (KDE) and Kullback-Leibler (KL) divergence to estimate the sufficient data volume for training authentication models. We introduce the Mouse Authentication Unit (MAU), leveraging Approximate Entropy (ApEn) to optimize segment length for efficient and accurate behavioral representation. Furthermore, we design the Local-Time Mouse Authentication (LT-AMouse) framework, integrating 1D-ResNet for local feature extraction and GRU for modeling long-term temporal dependencies. Taking the Balabit and DFL datasets as examples, we significantly reduced the data scale, particularly by a factor of 10 for the DFL dataset, greatly alleviating the training burden. Additionally, we determined the optimal input recognition unit length for the user authentication system on different datasets based on the slope of Approximate Entropy. Training with imbalanced samples, our model achieved a successful defense AUC 98.52% for blind attack on the DFL dataset and 94.65% on the Balabit dataset, surpassing the current sota performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。