ATLAS是跨学科高难度科学推理评测集,能区分顶尖大模型的真实推理能力。
ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
- 由专家原创800道跨学科题,杜绝数据泄露
- 要求多步推理与LaTeX表达,答案更接近真实科研
- 适合评估大模型在复杂科学任务中的通用推理能力
大型语言模型(LLMs)在现有基准上性能已趋于饱和,难以区分前沿模型。同时,现有高难度基准常存在领域单一、答案格式简单、易受数据污染等问题,与真实科学探究存在差距。为此,我们提出ATLAS(面向AGI的科学逻辑应用测试平台),一个大规模、高难度、跨学科的评估体系,包含约800道原创题目,覆盖数学、物理、化学、生物、计算机、地球科学和材料科学七个核心领域。其特点包括:(1)高度原创性与抗污染性,所有题目均为新创或深度改编;(2)跨学科融合设计,评估模型跨领域知识整合与推理能力;(3)高保真答案,强调多步推理与LaTeX格式表达,而非简单选择题;(4)严格质量控制,通过多阶段专家评审与对抗测试确保题目难度、科学价值与正确性。我们还提出使用多模型评委进行自动化、精细化的答案评估。初步实验表明,ATLAS能有效区分主流模型的高级科学推理能力。未来计划将其发展为长期、开放、社区驱动的平台,为通向人工智能通用智能提供可靠的“量尺”。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has led to performance saturation on many established benchmarks, questioning their ability to distinguish frontier models. Concurrently, existing high-difficulty benchmarks often suffer from narrow disciplinary focus, oversimplified answer formats, and vulnerability to data contamination, creating a fidelity gap with real-world scientific inquiry. To address these challenges, we introduce ATLAS (AGI-Oriented Testbed for Logical Application in Science), a large-scale, high-difficulty, and cross-disciplinary evaluation suite composed of approximately 800 original problems. Developed by domain experts (PhD-level and above), ATLAS spans seven core scientific fields: mathematics, physics, chemistry, biology, computer science, earth science, and materials science. Its key features include: (1) High Originality and Contamination Resistance, with all questions newly created or substantially adapted to prevent test data leakage; (2) Cross-Disciplinary Focus, designed to assess models' ability to integrate knowledge and reason across scientific domains; (3) High-Fidelity Answers, prioritizing complex, open-ended answers involving multi-step reasoning and LaTeX-formatted expressions over simple multiple-choice questions; and (4) Rigorous Quality Control, employing a multi-stage process of expert peer review and adversarial testing to ensure question difficulty, scientific value, and correctness. We also propose a robust evaluation paradigm using a panel of LLM judges for automated, nuanced assessment of complex answers. Preliminary results on leading models demonstrate ATLAS's effectiveness in differentiating their advanced scientific reasoning capabilities. We plan to develop ATLAS into a long-term, open, community-driven platform to provide a reliable "ruler" for progress toward Artificial General Intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。