arXiv:2505.24621cs.CL2025-05EMNLP被引 11

测试大模型破解加密文本能力,发现其存在安全漏洞风险。

Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities

  • 构建多领域加密文本数据集,评估大模型零样本与少样本解密能力。
  • 大模型在部分场景下可成功解密,但对侧信道攻击易受威胁。
  • 揭示大模型通用性缺陷,警示其在安全领域的双面风险。

大语言模型(LLMs)在自然语言理解与生成方面取得显著进展,但其在密码分析——这一关乎数据安全的关键领域——的评估仍不充分。为填补此空白,我们评估了前沿大模型对多种加密算法生成密文的破解能力。构建了一个涵盖多领域、多长度、多样写作风格和主题的明文-密文配对数据集。采用零样本、少样本及思维链提示方法,评估模型解密成功率并分析其理解能力。结果表明,大模型在特定场景下具备一定解密能力,但在侧信道攻击中暴露脆弱性,提示其可能因泛化不足而遭攻击。本研究凸显大模型在安全应用中的双重用途,并推动关于人工智能安全与风险的讨论。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have transformed natural language understanding and generation, leading to extensive benchmarking across diverse tasks. However, cryptanalysis - a critical area for data security and its connection to LLMs' generalization abilities - remains underexplored in LLM evaluations. To address this gap, we evaluate the cryptanalytic potential of state-of-the-art LLMs on ciphertexts produced by a range of cryptographic algorithms. We introduce a benchmark dataset of diverse plaintexts, spanning multiple domains, lengths, writing styles, and topics, paired with their encrypted versions. Using zero-shot and few-shot settings along with chain-of-thought prompting, we assess LLMs' decryption success rate and discuss their comprehension abilities. Our findings reveal key insights into LLMs' strengths and limitations in side-channel scenarios and raise concerns about their susceptibility to under-generalization-related attacks. This research highlights the dual-use nature of LLMs in security contexts and contributes to the ongoing discussion on AI safety and security.

密码分析大模型安全侧信道攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。