arXiv:2505.20658cs.CL2025-05ACL被引 15

用多样外部知识提升自然语言转时序逻辑的准确性

Enhancing Transformation from Natural Language to Signal Temporal Logic Using LLMs with Diverse External Knowledge

  • 基于聚类引导的LLM生成+规则过滤,构建1.6万样本多样化数据集
  • 新框架在两个数据集上准确率显著优于基线模型
  • 适合形式化验证、自动驾驶等需要精确语义描述的研究者

时序逻辑(TL),尤其是信号时序逻辑(STL),能实现精确的形式化规范,在自动驾驶和机器人等网络物理系统中广泛应用。自动将自然语言(NL)转换为STL是克服人工转换耗时易错问题的有力方法,但因缺乏数据集而进展有限。本文提出名为STL-DivEn的NL-STL数据集,包含16,000个经过多样化模式增强的样本。构建过程包括:先人工创建小规模种子集,再通过聚类识别代表性样本,引导大语言模型(LLMs)生成更多配对;最后通过严格的规则过滤和人工验证确保多样性和准确性。此外,提出知识引导的STL转换框架KGST,采用生成-精炼流程结合外部知识。统计分析显示,STL-DivEn比现有数据集更具多样性。指标评估与人工评价均表明,KGST在STL-DivEn和DeepSTL数据集上的转换准确率显著优于基线模型。

原文摘要 · Abstract (English)

Temporal Logic (TL), especially Signal Temporal Logic (STL), enables precise formal specification, making it widely used in cyber-physical systems such as autonomous driving and robotics. Automatically transforming NL into STL is an attractive approach to overcome the limitations of manual transformation, which is time-consuming and error-prone. However, due to the lack of datasets, automatic transformation currently faces significant challenges and has not been fully explored. In this paper, we propose an NL-STL dataset named STL-Diversity-Enhanced (STL-DivEn), which comprises 16,000 samples enriched with diverse patterns. To develop the dataset, we first manually create a small-scale seed set of NL-STL pairs. Next, representative examples are identified through clustering and used to guide large language models (LLMs) in generating additional NL-STL pairs. Finally, diversity and accuracy are ensured through rigorous rule-based filters and human validation. Furthermore, we introduce the Knowledge-Guided STL Transformation (KGST) framework, a novel approach for transforming natural language into STL, involving a generate-then-refine process based on external knowledge. Statistical analysis shows that the STL-DivEn dataset exhibits more diversity than the existing NL-STL dataset. Moreover, both metric-based and human evaluations indicate that our KGST approach outperforms baseline models in transformation accuracy on STL-DivEn and DeepSTL datasets.

时序逻辑自然语言大模型形式化验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。