arXiv:2511.20694cs.AIastro-ph.SR2025-11中稿 · NeurIPS

构建太阳物理推理数据集,测试大模型科学推理能力

Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning

  • 基于暑期学校题集构建结构化问答数据,含推理步骤与格式提示
  • 多智能体协作分解任务比单次提示在演绎推理上提升显著
  • 适合研究科学推理、多智能体系统与天体物理教育的学者

在太阳物理学领域,大语言模型的科学推理不仅涉及事实记忆,还需引入物理假设、保持单位一致,并以清晰的科学格式呈现。为此,我们提出《与恒星共思》数据集,适用于科学推理任务,并提供初步基准测试。数据源自美国国家航空航天局与大气研究中心联合举办的“与恒星共存”暑期学校习题,整理为包含问题背景、推理步骤、预期答案类型、真实答案、格式提示和元数据的问答结构。通过程序化评分器,采用考虑单位容差、符号等价性和模式验证的方法评估预测结果。我们对单次提示基线及四种多智能体模式进行基准测试,发现基于系统工程原则分解工作流,在需要演绎推理的问题上优于直接提示。

原文摘要 · Abstract (English)

Scientific reasoning through Large Language Models in heliophysics involves more than just recalling facts: it requires incorporating physical assumptions, maintaining consistent units, and providing clear scientific formats through coordinated approaches. To address these challenges, we present Reasoning With a Star, a newly contributed heliophysics dataset applicable to reasoning; we also provide an initial benchmarking approach. Our data are constructed from National Aeronautics and Space Administration & University Corporation for Atmospheric Research Living With a Star summer school problem sets and compiled into a readily consumable question-and-answer structure with question contexts, reasoning steps, expected answer type, ground-truth targets, format hints, and metadata. A programmatic grader checks the predictions using unit-aware numerical tolerance, symbolic equivalence, and schema validation. We benchmark a single-shot baseline and four multi-agent patterns, finding that decomposing workflows through systems engineering principles outperforms direct prompting on problems requiring deductive reasoning rather than pure inductive recall.

科学推理多智能体太阳物理数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。