首个面向密码学的超大规模问答数据集,助力大模型提升密码推理能力。
CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography
- 构建含200万+问答对的密码学专用数据集,覆盖学术文献与上下文元数据。
- 15个顶尖大模型在密码数学推理上表现不佳,平均准确率不足60%。
- 可训练定制化密码学助手,适合密码研究者与安全开发人员使用。
大型语言模型(LLMs)在通用自然语言任务中表现出色,但在密码学所需的深度推理与数学分析方面能力仍不明确,主要因缺乏合适的评估与训练数据。为此,我们提出CryptoQA,首个专为密码学设计的大规模问答数据集。该数据集包含超过两百万个来自精选学术来源的问答对,并附带可用于测试与训练的上下文元数据。我们在CryptoQA上基准测试了15个最先进的大模型,评估其事实准确性、数学推理、一致性、引用能力、逆向推理及对抗样本鲁棒性。除定量指标外,还提供专家评审以定性评估输出并建立黄金标准基准。结果表明,大模型在需形式化推理与精确数学知识的任务上存在显著性能缺陷,凸显亟需针对密码学研发的专用大模型助手。实验显示,利用CryptoQA微调后,大模型在密码任务上的表现明显提升。
原文摘要 · Abstract (English)
Large language models (LLMs) excel at many general-purpose natural language processing tasks. However, their ability to perform deep reasoning and mathematical analysis, particularly for complex tasks as required in cryptography, remains poorly understood, largely due to the lack of suitable data for evaluation and training. To address this gap, we present CryptoQA, the first large-scale question-answering (QA) dataset specifically designed for cryptography. CryptoQA contains over two million QA pairs drawn from curated academic sources, along with contextual metadata that can be used to test the cryptographic capabilities of LLMs and to train new LLMs on cryptographic tasks. We benchmark 15 state-of-the-art LLMs on CryptoQA, evaluating their factual accuracy, mathematical reasoning, consistency, referencing, backward reasoning, and robustness to adversarial samples. In addition to quantitative metrics, we provide expert reviews that qualitatively assess model outputs and establish a gold-standard baseline. Our results reveal significant performance deficits of LLMs, particularly on tasks that require formal reasoning and precise mathematical knowledge. This shows the urgent need for LLM assistants tailored to cryptography research and development. We demonstrate that, by using CryptoQA, LLMs can be fine-tuned to exhibit better performance on cryptographic tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。