arXiv:2606.19700cs.CL2026-06

用小模型自动提取火星改造文献知识,生成可直接使用的结构化数据。

TerraMARS: A Domain-Adapted Small-Language-Model Pipeline for Mars Terraforming Literature

论文配图:TerraMARS: A Domain-Adapted Small-Language-Model Pipeline for Mars Terraforming Literature
图 1 · 摘自论文原文
  • 用领域微调的小模型从科学文本中提取信息
  • 在开放论文语料上实现高精度问答与结构化输出
  • 适合做火星数字孪生和宜居性建模的研究者使用

研究人员致力于了解火星,以期未来使其适合人类居住。为此,需全面掌握该行星的大气、水文、表面化学、辐射环境及空间特征等科学文献中的信息。这些内容包含有价值的定量约束,可用于其他模型与研究,如宜居性评估与未来火星改造研究。本文提出TerraMARS,一个端到端的信息抽取流程,结合领域自适应的小语言模型,回答火星改造相关问题,并将非结构化火星科学文本转化为机器可读的JSON格式结构化输出。通过多阶段检索与分块框架处理开放获取论文语料库。使用量化低秩适配(QLoRA)对Google Gemma 3 1B进行领域微调,训练数据涵盖火星特定问答与信息抽取任务。该流程可生成两类输出,为数字孪生与火星宜居性建模等下游应用提供知识基础。当前输出表现良好,但仍需提升提取准确率与事实一致性。

原文摘要 · Abstract (English)

Researchers are interested in learning about Mars so that it may eventually become habitable for humans. To achieve this, there is a need for comprehensive knowledge of the planet's atmosphere, hydrology, surface chemistry, radiation environment, and spatial features through the scientific literature. These contain valuable information and meaningful quantitative constraints that can be used in other models and studies, such as habitability assessment and future terraforming studies. We present TerraMARS, an end-to-end information extraction pipeline that combines a domain-adapted Small Language Model to answer Mars terraforming-related questions and convert unstructured Mars science text into machine-readable structured outputs in JavaScript Object Notation (JSON) format. A corpus of open-access papers is collected and processed using a multistage retrieval and chunking framework. Google Gemma 3 1B was adapted to the domain using Quantized Low-Rank Adaptation (QLoRA) fine-tuning on Mars-specific question-answering and information extraction datasets. The resulting pipeline generates both types of output and provides a foundation for integrating knowledge from scientific literature into downstream applications like digital twins and habitability modeling for Mars. The output from this pipeline looks promising, but further improvements are needed to increase extraction accuracy and factual consistency.

小模型知识提取火星研究结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。