arXiv:2410.03569cs.LG2024-10被引 4

通过定制数据分布和损失函数,让机器学习更擅长模运算,破解更难的密码问题。

Making Hard Problems Easier with Custom Data Distributions and Loss Regularization: A Case Study in Modular Arithmetic

  • 用特定数据分布和结构化损失函数训练模型,提升模运算能力。
  • 可处理最多128个元素模q≤974269的运算,比之前强两倍。
  • 不仅适用于LWE,还能改进复制、关联回忆等经典任务,适合密码学与机器学习交叉研究者。

近期研究表明,基于机器学习的攻击在某些条件下优于传统的代数攻击,能针对后量子密码中的学习误差(LWE)问题发起有效攻击。然而,现有ML攻击难以扩展到更复杂的LWE设置。先前工作指出,这主要源于训练模型进行模运算的困难。为此,本文提出新方法:采用定制的训练数据分布和精心设计的损失函数,显著提升模型对模运算的学习能力,使其能够处理最多包含128个元素的模运算,模数q不超过974,269。我们将该方法应用于LWE问题,实现比以往工作高出两倍难度的密钥恢复。此外,该方法也提升了模型在复制、关联回忆和奇偶校验等经典问题上的表现,推动了跨领域研究。

原文摘要 · Abstract (English)

Recent work showed that ML-based attacks on Learning with Errors (LWE), a hard problem used in post-quantum cryptography, outperform classical algebraic attacks in certain settings. Although promising, ML attacks struggle to scale to more complex LWE settings. Prior work connected this issue to the difficulty of training ML models to do modular arithmetic, a core feature of the LWE problem. To address this, we develop techniques that significantly boost the performance of ML models on modular arithmetic tasks, enabling the models to sum up to $N=128$ elements modulo $q \le 974269$. Our core innovation is the use of custom training data distributions and a carefully designed loss function that better represents the problem structure. We apply an initial proof of concept of our techniques to LWE specifically and find that they allow recovery of 2x harder secrets than prior work. Our techniques also help ML models learn other well-studied problems better, including copy, associative recall, and parity, motivating further study.

机器学习密码学模运算深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。