arXiv:2608.05439cs.AIcs.LG2026-08

让AI学会拒绝生成不可靠的指令规范,提升机器人安全决策能力。

SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications

论文配图:SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications
图 1 · 摘自论文原文
  • 通过双重信号评估翻译可靠性:自然语言回译一致性和重复翻译差异性。
  • 在三种逻辑语言上实现90%以上正确率,且能主动拒绝不可信输入。
  • 适合需要高可靠性的自动驾驶、机器人控制等安全关键场景。

将自然语言指令转化为机器可读的形式化规范,使机器人和自主系统能够规划、推理并形式化验证其行为。然而,现有翻译模型通常对每个输入都生成规范,即使结果不可靠或未能捕捉用户意图,在安全关键应用中带来风险。受选择性置信预测启发,我们提出一种选择性翻译框架,不仅能生成形式化规范,还能判断何时可信任。可靠性通过两个互补的黑盒信号评分:规范回译为自然语言的保真度,以及在语义等价下重复翻译的分散程度,二者对不同错误敏感,联合分离错误翻译更有效。置信风险控制将该评分转化为接受或放弃规范的决策,保证错误规范被采纳率有分布无关的上限;同时,基于指令嵌入的置信异常检测器可在翻译前筛除分布外输入。该框架通用性强,适用于多种形式规范语言,实验在信号时序逻辑(STL)、线性时序逻辑(LTL)及几何时空逻辑(SpaTiaL)上均证明了更高的翻译可靠性、跨层级扰动下的鲁棒性,以及有效的不确定性感知弃权能力。本工作为可信自然语言接口奠定基础,使AI系统能识别生成规范不可靠的情况。

原文摘要 · Abstract (English)

Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, reason, and formally verify their behavior. However, existing translation models typically generate a specification for every input, even when the result is unreliable or fails to capture the user's intent, creating risks in safety-critical applications. Inspired by selective conformal prediction, we propose a selective translation framework that not only generates formal specifications but also determines when they can be trusted. Reliability is scored by two complementary black-box signals, the fidelity of the specification back-translated into natural language and the dispersion of repeated translations under exact semantic equivalence, which fail on different errors and jointly separate incorrect translations more sharply than either alone. Conformal risk control calibrates this score into a decision that accepts a specification or abstains, with a distribution-free bound on the rate at which incorrect specifications are accepted for execution, and a conformal anomaly detector on instruction embeddings screens out-of-distribution inputs before any translation is attempted. The proposed framework is general across formal specification languages, with experiments on Signal Temporal Logic (STL), Linear Temporal Logic (LTL), and geometric Spatio-Temporal Logic (SpaTiaL) demonstrating improved translation reliability, robustness under the evaluated cross-tier shifts, and effective uncertainty-aware abstention. This work establishes a foundation for trustworthy natural language interfaces by enabling AI systems to recognize when generated specifications may not be reliable.

自然语言形式化验证不确定性可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。