评估法律本体学习中语义是否丢失,发现不同模型方法组合差异巨大。
On Measuring Semantic Preservation in Legal Ontology Learning

- 用大模型在原文和结构化表示上的表现差值衡量语义损失
- 在并购协议分析中发现语义损失随推理复杂度显著变化
- 为法律知识系统选型提供实证依据,适合法律AI研究者
本体学习将非结构化文本转化为可用于自动推理的结构化表示,但这一过程可能造成语义丢失。现有评估方法只关注结构正确性,无法检测语义保留情况。本文提出一种新评估方法:比较大语言模型(LLM)在原始文档与转换后表示上的任务表现,差异即为语义损失量。我们在法律并购协议分析领域验证该方法,该领域语言复杂且语义要求精确。对比了直接使用LLM与三种本体学习方法在六种语言模型上的表现。结果表明,语义损失存在系统性,且受推理复杂度和模型-方法交互影响显著。贡献包括:(1) 提出衡量本体学习中语义保留的评估框架;(2) 实证发现语义损失因模型-方法组合而异,为法律知识系统配置提供指导。
原文摘要 · Abstract (English)
Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on structural correctness while failing to measure whether meaning survives transformation. We propose an evaluation methodology that addresses this: comparing LLM task performance on source documents against performance on transformed representations, with the difference quantifying semantic loss. We demonstrate this approach on legal merger agreement analysis, a domain chosen for its complex language and precise semantic requirements, comparing direct LLM application against three ontology learning methods across six language models. The results reveal systematic semantic loss with significant variation based on reasoning complexity and model-method interactions. Our contributions are: (1) an evaluation framework for measuring semantic preservation in ontology learning, and (2) empirical evidence that semantic loss varies dramatically with model-method pairing, providing guidance for selecting optimal configurations in legal knowledge systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。