用统一框架构建企业级语义知识库,提升复杂问答能力
Construct, Align, and Reason: Large Ontology Models for Enterprise Knowledge Management
- 分层构建结构化与非结构化数据融合的语义知识图谱
- 40亿参数模型在复杂推理任务上达89.47%准确率,优于DeepSeek-V3.2
- 适合需要精准知识推理的企业级应用,如智能客服、决策支持
企业级知识管理面临多源异构数据整合与有效语义推理的挑战。传统知识图谱在隐式关系发现和复杂问答理解方面存在不足。为此,我们提出统一的构建-对齐-推理框架,即大语义模型(LOM)。首先从结构化数据库与非结构化文本中构建双层企业语义本体,并融合为完整企业语义本体。为实现指令对齐推理,提出三阶段训练流程:本体指令微调以增强结构理解;文本-本体对齐以强化节点语义编码;基于课程学习的本体-语言对多任务指令微调以提升语义推理与生成能力。同时构建涵盖多样本体推理任务的训练与评估数据集。在该基准上,40亿参数的LOM达到89.47%准确率,显著优于DeepSeek-V3.2在复杂图推理任务上的表现,证明了本体结构与语言信息的有效融合。
原文摘要 · Abstract (English)
Enterprise-scale knowledge management faces significant challenges in integrating multi-source heterogeneous data and enabling effective semantic reasoning. Traditional knowledge graphs often struggle with implicit relationship discovery and lack sufficient semantic understanding for complex question answering. To address these limitations, we introduce a unified construct--align--reason framework, the large ontology model (LOM). We first build a dual-layer enterprise ontology from structured databases and unstructured text, subsequently fusing these sources into a comprehensive enterprise ontology. To enable instruction-aligned reasoning, we propose a unified three-stage training pipeline: ontology instruction fine-tuning to improve structural understanding; text-ontology grounding to strengthen node semantic encoding; and multi-task instruction tuning on ontology-language pairs with curriculum learning to enhance semantic reasoning and generation. We also construct comprehensive training and evaluation datasets covering diverse ontology reasoning tasks. On this benchmark, our 4B-parameter LOM achieves 89.47% accuracy and outperforms DeepSeek-V3.2 on complex graph reasoning, indicating effective fusion of ontology structure and language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。