用AI和专家知识构建可追溯的机场运营知识图谱
Semi-Automated Knowledge Engineering and Process Mapping for Total Airport Management
- 结合专家知识与大模型,分两阶段融合生成语义一致的知识三元组
- 文档级处理比片段推理更有效捕捉非线性流程依赖关系
- 确保每条信息可溯源,适合需要高可信度的航空运维场景
机场运营文档因专业术语繁多、法规严格、区域信息私有及多方沟通碎片化,导致数据孤岛与语义不一致,严重阻碍全面机场管理(TAM)推进。本文提出一种方法框架,通过符号知识工程(KE)与生成式大语言模型(LLM)双阶段融合,构建领域对齐、机器可读的知识图谱(KG)。框架采用引导式融合策略,由专家构建的KE结构指导LLM提示词,提升语义对齐的知识三元组发现效率。在Google LangExtract库上评估该方法,并比较局部片段推理与文档级处理在上下文窗口利用上的差异。与以往长上下文退化的观察相反,文档级处理显著提升非线性流程依赖的恢复能力。为保障机场运营所需的高保真溯源性,框架融合概率模型用于知识发现,确定性算法用于将每项抽取结果锚定至原始来源,实现绝对可追溯与可验证。最终,引入自动化框架,将该流程流水线化,从非结构化文本语料中合成复杂操作流程。
原文摘要 · Abstract (English)
Documentation of airport operations is inherently complex due to extensive technical terminology, rigorous regulations, proprietary regional information, and fragmented communication across multiple stakeholders. The resulting data silos and semantic inconsistencies present a significant impediment to the Total Airport Management (TAM) initiative. This paper presents a methodological framework for constructing a domain-grounded, machine-readable Knowledge Graph (KG) through a dual-stage fusion of symbolic Knowledge Engineering (KE) and generative Large Language Models (LLMs). The framework employs a scaffolded fusion strategy in which expert-curated KE structures guide LLM prompts to facilitate the discovery of semantically aligned knowledge triples. We evaluate this methodology on the Google LangExtract library and investigate the impact of context window utilization by comparing localized segment-based inference with document-level processing. Contrary to prior empirical observations of long-context degradation in LLMs, document-level processing improves the recovery of non-linear procedural dependencies. To ensure the high-fidelity provenance required in airport operations, the proposed framework fuses a probabilistic model for discovery and a deterministic algorithm for anchoring every extraction to its ground source. This ensures absolute traceability and verifiability, bridging the gap between "black-box" generative outputs and the transparency required for operational tooling. Finally, we introduce an automated framework that operationalizes this pipeline to synthesize complex operational workflows from unstructured textual corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。