arXiv:2603.09231cs.AI2026-03

用认知分层方法生成高质量数据,让大模型更好理解太空态势感知。

Cognitively Layered Data Synthesis for Domain Adaptation of LLMs to Space Situational Awareness

  • 按认知层级设计问题,覆盖从记忆到创造的九类题目
  • 构建23万条数据,使模型在专业任务上表现提升超170%
  • 适合需要高精度工程推理的大模型落地场景

大语言模型在通用任务中表现优异,但在太空态势感知(SSA)等复杂工程领域迁移仍面临挑战,主要源于任务链结构对齐不足、高级认知监督缺失以及数据质量标准与工程需求不匹配。核心瓶颈在于高质量监督微调(SFT)数据集的构建。为此,我们提出基于布卢姆分类学的领域特定微调数据生成框架BD-FDG,通过三种机制解决知识覆盖不全、认知深度浅和质量控制弱的问题:采用知识树确保语料结构化覆盖,设计涵盖九类主题和六级认知层次(从记忆到创造)的问题生成方案以实现难度连续梯度,建立多维评分流水线保障领域严谨性与一致性。利用该框架,我们构建了约23万样本的SSA-SFT数据集,并微调Qwen3-8B得到SSA-LLM-8B。实验表明,该模型在领域测试集上相对BLEU-1得分分别提升144%(无思考)和176%(有思考),在对战评测中胜率高达82.21%,同时在MMLU-Pro、MATH-500等通用基准上性能基本保持不变。结果验证了基于认知分层的数据构造在复杂工程领域的有效性,提供了可复用的领域适配框架。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate exceptional performance on general-purpose tasks. however, transferring them to complex engineering domains such as space situational awareness (SSA) remains challenging owing to insufficient structural alignment with mission chains, the absence of higher-order cognitive supervision, and poor correspondence between data quality criteria and engineering specifications. The core bottleneck is the construction of high-quality supervised fine-tuning (SFT) datasets. To this end, we propose BD-FDG (Bloom's Taxonomy-based Domain-specific Fine-tuning Data Generation), a framework that addresses incomplete knowledge coverage, shallow cognitive depth, and limited quality controllability through three mechanisms: structured knowledge organization, cognitively layered question modeling, and automated quality control. The framework uses a knowledge tree to ensure structured corpus coverage, designs a question generation scheme spanning nine categories and six cognitive levels from Remember to Create to produce samples with a continuous difficulty gradient, and applies a multidimensional scoring pipeline to enforce domain rigor and consistency. Using BD-FDG, we construct SSA-SFT, a domain dataset of approximately 230K samples, and fine-tune Qwen3-8B to obtain SSA-LLM-8B. Experiments show that SSA-LLM-8B achieves relative BLEU-1 improvements of 144\% (no-think) and 176\% (think) on the domain test set and a win rate of 82.21\% over the baseline in arena comparisons, while largely preserving general benchmark performance (MMLU-Pro, MATH-500). These results validate SFT data construction driven by cognitive layering as an effective paradigm for complex engineering domains and provide a transferable framework for domain-specific LLM adaptation.

大模型适配认知分层数据合成太空感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。