arXiv:2604.11778cs.CLcs.AI2026-04被引 6

构建365个通用推理任务,评估大模型跨领域推理能力。

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks

论文配图:General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
图 1 · 摘自论文原文
  • 限定知识为中小学水平,剥离专业背景干扰
  • 26个模型最高仅达62.8%准确率,远低于数学物理表现
  • 适合关注大模型泛化推理能力的研究者

当前大型语言模型在数学、物理等特定领域展现出卓越的推理能力,但在更广泛、更具挑战性的通用推理方面仍缺乏深入研究。与依赖专业知识的领域推理不同,通用推理虽不依赖专家知识,却面临复杂约束、嵌套逻辑分支和语义干扰等难题。为此,我们提出General365基准,专为评估大模型的通用推理能力而设计。该基准将背景知识限制在K-12水平,明确剥离推理与专业知识的关联。包含365个基础问题及1,095个变体问题,覆盖八大类别,确保高难度与多样性。对26个主流大模型的评估显示,即使表现最优的模型也仅达到62.8%准确率,显著低于其在数学和物理基准上的接近满分表现。结果表明,当前大模型的推理能力高度依赖具体领域,向更广泛应用场景迁移仍有巨大提升空间。我们期望General365能推动大模型推理从专用任务迈向真正通用的现实场景。代码、数据集与排行榜:https://general365.github.io

原文摘要 · Abstract (English)

Contemporary large language models (LLMs) have demonstrated remarkable reasoning capabilities, particularly in specialized domains like mathematics and physics. However, their ability to generalize these reasoning skills to more general and broader contexts--often termed general reasoning--remains under-explored. Unlike domain-specific reasoning, general reasoning relies less on expert knowledge but still presents formidable reasoning challenges, such as complex constraints, nested logical branches, and semantic interference. To address this gap, we introduce General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning from specialized expertise. The benchmark comprises 365 seed problems and 1,095 variant problems across eight categories, ensuring both high difficulty and diversity. Evaluations across 26 leading LLMs reveal that even the top-performing model achieves only 62.8% accuracy, in stark contrast to the near-perfect performances of LLMs in math and physics benchmarks. These results suggest that the reasoning abilities of current LLMs are heavily domain-dependent, leaving significant room for improvement in broader applications. We envision General365 as a catalyst for advancing LLM reasoning beyond domain-specific tasks toward robust, general-purpose real-world scenarios. Code, Dataset, and Leaderboard: https://general365.github.io

通用推理大模型评测基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。