构建航天器行为推理的高保真基准,助力太空态势感知
AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

- 基于真实轨道仿真与观测数据设计三类可验证推理任务
- 多模型测试显示无单一模型全面领先,大小模型各有所长
- 强调物理约束与战术理解并重,适合航天智能研究者使用
理解航天器为何机动(而不仅是是否机动)是日益重要的太空态势感知问题。现有分析流程侧重检测,难以深入推断行为含义。AstroMind是一个基于物理规律的基准,融合高保真轨道动力学模拟与真实观测约束,转化为三类可验证推理任务:意图推断、机动参数估计和威胁评估。每个场景包含真实传感噪声及多源文本情报,可靠性各异。评估指标同时考量语义正确性与物理约束下的定量一致性。对一系列开源模型的测试表明,无模型在所有维度占优:Qwen3(32B)在意图推断准确率上领先;QwQ(32B)在威胁评估和解析项相对误差最小;GPT-OSS(20B)推理质量最佳,提取了136/241个标量参数。训练数据构成与推理风格与模型规模同等重要。结构化提示在8B模型中均有提升,尤其对已能追踪物理约束的模型增益更大。AstroMind为需要兼顾物理正确性与战术理解的问题提供了统一评测标准。
原文摘要 · Abstract (English)
Understanding why a spacecraft maneuvers -- rather than simply that it did -- is an increasingly important problem for space domain awareness as Earth orbits grow crowded and contested. Current analysis pipelines are built for detection: they are good at picking up that something happened, less good at reasoning about what it means. AstroMind is a physics-grounded benchmark designed to close that gap. It draws on high-fidelity astrodynamics simulations and real observational constraints, converting them into verifiable reasoning problems across three task types: intent inference, maneuver parameter estimation, and threat assessment. Each scenario includes realistic sensing noise and multi-source textual intelligence at varying reliability levels. Evaluation metrics capture both semantic correctness and quantitative consistency under physical constraints. Benchmarking a suite of open-weight models shows no single model dominates every axis: Qwen3 (32B) leads on intent inference accuracy; QwQ (32B) leads on threat assessment and achieves the lowest median relative error on parsed items; GPT-OSS (20B) produces the strongest judged reasoning quality and extracts the most scalar values for parameter estimation (136 of 241 parsed items). Training data composition and reasoning style matter as much as model size. Structured reasoning prompts help consistently across tested 8B models, with larger gains for those that can already track physical constraints. AstroMind gives the field a shared test for a problem where getting the physics right and reading the tactical situation correctly are both required -- neither is sufficient on its own.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。