将火星探测任务文档从自然语言转为逻辑表达,助力机器人自主决策。
A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration

- 基于真实任务文档构建自然语言到一阶逻辑的转换数据集。
- 涵盖2003至2013年多个任务阶段,包含时间结构与操作依赖关系。
- 适合研究形式化推理与语言理解交叉方向的学者使用。
未来行星探测设想在通信受限、无全球定位且人类干预极少的环境下,由自主机器人执行任务。在此类环境中,机器人不仅需感知与行动,还需对任务目标、操作约束和环境变化进行推理。尽管已有研究集中于感知与控制,但如何将高层任务知识转化为结构化、机器可读的表示仍缺乏探索。本文提出一个面向行星探测领域的自然语言(NL)到一阶逻辑(FOL)翻译的试点基准。数据集源自美国国家航空航天局行星数据系统(PDS)的真实任务文档,覆盖2003至2013年间多个任务阶段,包括发射、加速、巡航、滑行及轨道运行等,以丰富自然语言描述。我们手动标注了对应的一阶逻辑表示,捕捉时间结构、代理角色与操作依赖。此外,提供结构化的谓词词汇表与类型化常量,支持在不同先验知识水平下进行可控实验。该试点基准为语言理解与形式推理交叉研究提供了基于真实、高安全性任务数据的基础。数据集地址:https://github.com/HaydenMM/planetary-logic-benchmark/blob/main/pilot_benchmark.json
原文摘要 · Abstract (English)
Future planetary exploration envisions autonomous robotic agents operating under severe communication constraints, without global positioning, and with minimal human intervention. In such environments, agents must not only perceive and act, but also reason over mission objectives, operational constraints, and evolving environmental conditions. While prior work has largely focused on perception and control, the translation of high-level mission knowledge into structured, machine-interpretable representations remains underexplored. We introduce a pilot benchmark for translating natural language (NL) into First-Order Logic (FOL) within the domain of planetary exploration. The dataset is constructed from real mission documentation sourced from NASA's Planetary Data System (PDS), spanning missions from 2003 to 2013. These documents describe mission phases such as launch, boost, coast, cruise, and orbital operations in rich natural language. We manually annotate these documents with corresponding FOL representations that capture temporal structure, agent roles, and operational dependencies. In addition, we provide structured predicate vocabularies and typed constants to enable controlled experimentation with varying levels of prior knowledge. This pilot benchmark provides a foundation for research at the intersection of language understanding and formal reasoning, grounded in real-world, safety-critical mission data. The dataset is provided at: https://github.com/HaydenMM/planetary-logic-benchmark/blob/main/pilot_benchmark.json
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。