arXiv:2510.12080cs.AI2025-10被引 2

测试大模型生成随机数能力,发现效果不稳定且偏差明显。

Evaluating the Quality of Randomness and Entropy in Tasks Supported by Large Language Models

  • 设计多类随机任务实验,评估模型生成随机性表现。
  • 模型输出熵值偏低,NIST测试通过率不足40%。
  • 适合关注AI安全与可信随机生成的开发者参考。

大语言模型(LLM)在博弈、调度、智能体和密码学等需随机性的任务中应用广泛,但其生成与使用随机数的能力尚不明确。本文通过一系列实验,考察了外部工具可用性、任务类型、模型状态(新初始化与非新)及提示策略对随机任务表现的影响。实验涵盖随机数生成、密码类随机字符串生成、元素打乱及基于熵和NIST随机性测试套件的质量评估。结果表明,尽管模型输出表现出一定随机性,但性能极不稳定,常显著偏离预期行为。分析显示,现有模型在随机性生成上存在明显局限,亟需改进以有效支持随机性相关任务。

原文摘要 · Abstract (English)

The rapid advancement of large language model (LLM) technology has led to diverse applications, many of which inherently require randomness, such as stochastic decision-making, gaming, scheduling, AI agents, and cryptography-related tasks. However, the capabilities of LLMs in handling randomness, particularly in generating and utilizing random numbers effectively, remain unclear. This paper investigates the capacity of LLMs for handling tasks that involve randomness through a series of experiments. We designed a set of experiments that consider various factors that can influence an LLM's performance in tasks involving randomness, such as accessibility to external tools, types of tasks, model states (fresh vs. non-fresh), and prompting strategies. The experiments cover a range of tasks, including generating random numbers, generating random strings such as passwords, shuffling items, and evaluating the quality of randomness using entropy and the NIST randomness test-suite. Our findings reveal that while LLMs can generate outputs that exhibit some degree of randomness, their performance is inconsistent and often deviates significantly from the expected behavior. The analysis of the experimental results highlights key limitations and areas where improvement is needed for the LLMs to effectively handle tasks involving randomness

大模型随机性安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。