arXiv:2501.07600cs.LGcs.CV2025-01

研究键盘输入数据的广度与深度如何影响孪生网络性能

Impact of Data Breadth and Depth on Performance of Siamese Neural Network Model: Experiments with Three Keystroke Dynamic Datasets

  • 用特征空间和密度分析数据广度(人数)与深度(样本量)影响
  • 增加用户数能更好捕捉个体差异,自由文本数据对深度更敏感
  • 适合行为生物识别系统设计者参考,尤其关注数据构建策略

深度学习模型如孪生神经网络(SNN)在捕捉行为数据复杂模式方面展现出巨大潜力。然而,数据集广度(即用户数量)和深度(如每位用户的训练样本量)对模型性能的影响常被默认且缺乏深入探索。为此,我们在三个公开键盘输入数据集(Aalto、CMU、Clarkson II)上进行了大量实验,借助“特征空间”与“密度”概念,系统考察了训练用户数、每用户样本数、样本内数据量及训练三元组数量的影响。结果表明:在可行条件下,增加数据广度有助于训练出能有效捕捉跨用户差异的模型;而深度的影响因数据集类型而异:自由文本数据受样本数、序列长度、三元组数量和画廊样本大小三因素共同影响,样本不足易导致欠训练;固定文本数据则受这些因素影响较小,更易训练出优质模型。研究揭示了数据广度与深度在行为生物识别中的关键作用,为构建更有效的认证系统提供实践指导。

原文摘要 · Abstract (English)

Deep learning models, such as the Siamese Neural Networks (SNN), have shown great potential in capturing the intricate patterns in behavioral data. However, the impacts of dataset breadth (i.e., the number of subjects) and depth (e.g., the amount of training samples per subject) on the performance of these models is often informally assumed, and remains under-explored. To this end, we have conducted extensive experiments using the concepts of "feature space" and "density" to guide and gain deeper understanding on the impact of dataset breadth and depth on three publicly available keystroke datasets (Aalto, CMU and Clarkson II). Through varying the number of training subjects, number of samples per subject, amount of data in each sample, and number of triplets used in training, we found that when feasible, increasing dataset breadth enables the training of a well-trained model that effectively captures more inter-subject variability. In contrast, we find that the extent of depth's impact from a dataset depends on the nature of the dataset. Free-text datasets are influenced by all three depth-wise factors; inadequate samples per subject, sequence length, training triplets and gallery sample size, which may all lead to an under-trained model. Fixed-text datasets are less affected by these factors, and as such make it easier to create a well-trained model. These findings shed light on the importance of dataset breadth and depth in training deep learning models for behavioral biometrics and provide valuable insights for designing more effective authentication systems.

行为生物识别孪生网络键盘动态数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。