arXiv:2506.10585cs.AIcs.HC2025-06

用新数列测试大模型的符号推理能力,看它能否发现隐藏规律并正确推导结果。

Primender Sequence: A Novel Mathematical Construct for Testing Symbolic Inference and AI Reasoning

  • 定义新数列:包含质数或末位为质数的数,规则简单但结构复杂。
  • 测试多个大模型,发现只有少数能正确推导出10万项数列,准确率差异显著。
  • 适合研究大模型逻辑推理、数学规律识别和可解释性评估的科研人员使用。

本文提出一种名为Primender的新整数序列,其规则结合了经典素数与基于模运算的数字后缀条件:若一个数是素数,或其个位数字构成的任意长度后缀中至少有一个是素数,则该数属于该序列。此序列具有确定性但非平凡的结构,融合数论性质与符号模式特征。研究将该序列作为评估大语言模型(LLMs)符号推理能力的基准,旨在检验模型对隐藏规则的推断、数学假说的验证及符号逻辑的规模化泛化能力。核心假设为:当序列中某数恰好比不大于它的最大素数大1时,该数与前一项之差也为1。设计了结构化提示与评估框架,测试包括ChatGPT、Copilot、DeepSeek、Gemini、Grok和LLaMA在内的多款先进模型,任务包括识别规则、验证假设、生成后续10万项。通过规则推断准确率、假设验证、序列有效性与符号解释质量等指标进行对比分析。本工作贡献了一个新颖的数学构造与可复现的评测方法,推动数论、人工智能与软件工程的交叉研究。

原文摘要 · Abstract (English)

This paper introduces the Primender sequence, a novel integer sequence defined by a hybrid rule that combines classical primality with modular digit-based conditions. Specifically, a number n is included in the sequence if it is prime or ends with a prime number of unit digit or any length. In other words, numbers which are primes or have at least one prime suffix. The resulting sequence exhibits a deterministic yet non-trivial structure, blending number-theoretic properties with symbolic patterning. We propose the Primender sequence as a benchmark for evaluating the symbolic reasoning capabilities of Large Language Models (LLMs). The study is motivated by the need for interpretable, rule-based testbeds that can assess an LLM's ability to infer hidden rules, validate mathematical hypotheses, and generalize symbolic logic at scale. A key hypothesis explored is: Whenever a number in the Primender sequence is exactly one more than the largest prime less than or equal to it, the difference between it and the previous number in the sequence is also 1. We design a structured prompt and evaluation framework to test this hypothesis across multiple state-of-the-art LLMs, including ChatGPT, Copilot, DeepSeek, Gemini, Grok, and LLaMA. The models are tasked with identifying the underlying rule, validating the hypothesis, and generating the next 100,000 terms of the sequence. Comparative metrics such as rule inference accuracy, hypothesis evaluation, sequence validity, and symbolic explanation quality are used to assess model performance. This work contributes a novel mathematical construct and a reproducible methodology for benchmarking LLMs in symbolic reasoning, hypothesis testing, and scalable pattern generalization - bridging the domains of number theory, artificial intelligence, and software engineering.

符号推理大模型评测数列生成数学逻辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。