用大模型从论文中快速提取月球资源数据,提升探月规划效率
Towards Large Language Models for Lunar Mission Planning and In Situ Resource Utilization
- 用现成大模型解析科学论文中的表格数据,自动提取月球成分信息
- 模型对常见数据提取准确率高,但精细矿物分析仍有提升空间
- 适合需要快速获取月球资源分布的航天规划与科研人员
月球任务规划的关键在于评估当地原材料的可用性。然而,大量相关测量数据分散在各类科学文献中。本文探讨利用大语言模型(LLMs)快速处理科学文献语料库,以获取月球成分数据的可行性。尽管利用大模型从科学文档中提取知识并非新方法,但该应用面临独特挑战:月球样本异质性强,表征细节复杂,材料性质对成分微小变化敏感,因此准确性和不确定性量化尤为重要。研究发现,现成的LLMs在提取文献中常见表格数据方面总体有效。但当前方法仍有改进空间,尤其在捕捉精细矿物学信息及处理更复杂、微妙的信息方面。未来可进一步优化数据提取精度。
原文摘要 · Abstract (English)
A key factor for lunar mission planning is the ability to assess the local availability of raw materials. However, many potentially relevant measurements are scattered across a variety of scientific publications. In this paper we consider the viability of obtaining lunar composition data by leveraging LLMs to rapidly process a corpus of scientific publications. While leveraging LLMs to obtain knowledge from scientific documents is not new, this particular application presents interesting challenges due to the heterogeneity of lunar samples and the nuances involved in their characterization. Accuracy and uncertainty quantification are particularly crucial since many materials properties can be sensitive to small variations in composition. Our findings indicate that off-the-shelf LLMs are generally effective at extracting data from tables commonly found in these documents. However, there remains opportunity to further refine the data we extract in this initial approach; in particular, to capture fine-grained mineralogy information and to improve performance on more subtle/complex pieces of information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。