arXiv:2504.19093cs.CRcs.AI2025-04ACL被引 16

用密码学挑战测试大模型推理能力,发现现有模型仍有明显短板。

CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges

  • 构建2358个密码解密任务,覆盖5大领域14子领域
  • 顶尖模型在古典密码破解上表现不佳,存在明显能力缺口
  • 适合研究大模型推理局限与密码安全评估的学者参考

大型语言模型(LLMs)在数学和编程等领域的推理能力取得显著进展,如o1和o3系列模型。然而,在需要密码学专业知识的领域,其推理能力仍鲜有研究。本文提出CipherBank,一个全面的基准测试集,用于评估LLMs在密码解密任务中的推理能力。CipherBank包含2,358个精心设计的问题,涵盖5个领域、14个子领域及262种唯一明文,聚焦隐私敏感且贴近真实场景的加密任务。从密码学角度,该数据集涵盖3大类加密方法,涉及9种不同算法,从古典密码到定制加密技术均有覆盖。我们在GPT-4o、DeepSeek-V3以及前沿推理型模型o1和DeepSeek-R1上进行评估。结果表明,通用聊天模型与推理型模型之间、乃至当前推理型模型自身在古典密码解密任务中均存在显著性能差距,揭示了模型在理解与操作加密数据方面的挑战。通过详细分析与错误归因,我们总结出若干关键观察,为提升大模型在密码推理方面的表现提供了方向。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities, especially the recent advancements in reasoning, such as o1 and o3, pushing the boundaries of AI. Despite these impressive achievements in mathematics and coding, the reasoning abilities of LLMs in domains requiring cryptographic expertise remain underexplored. In this paper, we introduce CipherBank, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs in cryptographic decryption tasks. CipherBank comprises 2,358 meticulously crafted problems, covering 262 unique plaintexts across 5 domains and 14 subdomains, with a focus on privacy-sensitive and real-world scenarios that necessitate encryption. From a cryptographic perspective, CipherBank incorporates 3 major categories of encryption methods, spanning 9 distinct algorithms, ranging from classical ciphers to custom cryptographic techniques. We evaluate state-of-the-art LLMs on CipherBank, e.g., GPT-4o, DeepSeek-V3, and cutting-edge reasoning-focused models such as o1 and DeepSeek-R1. Our results reveal significant gaps in reasoning abilities not only between general-purpose chat LLMs and reasoning-focused LLMs but also in the performance of current reasoning-focused models when applied to classical cryptographic decryption tasks, highlighting the challenges these models face in understanding and manipulating encrypted data. Through detailed analysis and error investigations, we provide several key observations that shed light on the limitations and potential improvement areas for LLMs in cryptographic reasoning. These findings underscore the need for continuous advancements in LLM reasoning capabilities.

密码学大模型推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。