arXiv:2603.22136cs.CLcs.DB2026-03

构建分层语义框架,逐步将自然语言转化为机器可理解的知识。

The Semantic Ladder: A Framework for Progressive Formalization of Natural Language Content for Knowledge Graphs and AI Systems

  • 通过分层模块化结构,逐步提升自然语言的语义明确性。
  • 支持从文本到逻辑模型的连续转换,保留语义可追溯性。
  • 适合需融合多源异构数据的AI知识系统建设者使用。

语义数据与知识基础设施必须调和两种根本不同的表征形式:自然语言(绝大多数知识以该形式产生和传播)与形式化语义模型(支持机器可操作集成、互操作与推理)。在数据输入时实现完全语义形式化仍是一大核心挑战。本文提出语义阶梯框架,支持数据与知识的渐进式形式化。基于可识别的语义单元作为意义载体,该框架在语义显式度递增的多个层次间组织表示,涵盖自然语言片段、基于本体的模型及高阶逻辑模型。各层次间的转换支持语义增强、陈述结构化与逻辑建模,同时保持语义连续性与可追溯性。该方法实现语义知识空间的增量构建,减轻语义解析负担,并支持异构表示的整合,包括自然语言、结构化语义模型与向量嵌入。语义阶梯为可扩展、可互操作且适配AI的数据与知识基础设施提供了基础。

原文摘要 · Abstract (English)

Semantic data and knowledge infrastructures must reconcile two fundamentally different forms of representation: natural language, in which most knowledge is created and communicated, and formal semantic models, which enable machine-actionable integration, interoperability, and reasoning. Bridging this gap remains a central challenge, particularly when full semantic formalization is required at the point of data entry. Here, we introduce the Semantic Ladder, an architectural framework that enables the progressive formalization of data and knowledge. Building on the concept of modular semantic units as identifiable carriers of meaning, the framework organizes representations across levels of increasing semantic explicitness, ranging from natural language text snippets to ontology-based and higher-order logical models. Transformations between levels support semantic enrichment, statement structuring, and logical modelling while preserving semantic continuity and traceability. This approach enables the incremental construction of semantic knowledge spaces, reduces the semantic parsing burden, and supports the integration of heterogeneous representations, including natural language, structured semantic models, and vector-based embeddings. The Semantic Ladder thereby provides a foundation for scalable, interoperable, and AI-ready data and knowledge infrastructures.

知识图谱语义建模自然语言处理知识工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。