用案例推理对比法律形式化,发现大模型生成结果差异大且类型多样。
By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode
- 通过节点匹配和SAT求解器找不同形式化之间的分歧案例。
- 九个大模型对十条欧盟法规的形式化差异显著,结构相似但行为迥异。
- 生成的反例可被法律专家评估,揭示真实法律争议点。
法律条文的形式化有望实现机器可读法律与自动化法律推理,近期大语言模型使直接从法条文本生成形式化变得诱人。然而,任何形式化都隐含解释选择,其后果难以预判,尤其是当大模型是作者时。本文提出一种系统方法:通过分析不同形式化在具体案件中的推论差异来比较它们。给定同一法律条文的多个形式化,我们进行节点级匹配,为每对形式化构建共享接口,并利用SAT求解器枚举出两者不一致的边缘案例。这些案例被转化为具体的事实场景,供法律专家审查与处理。我们将该方法应用于九个前沿大模型对十条欧盟法律条文生成的形式化。结果显示,不同形式化间的推论行为差异与结构一致性基本无关;所生成的反例揭示了多种质性不同的分歧类型,包括与法律评论中真实争议相呼应的分歧。
原文摘要 · Abstract (English)
Formalizing legal provisions promises machine-accessible law and automated legal reasoning, and recent LLMs make it tempting to generate such formalizations directly from statutory text. However, any formalization makes implicit interpretive choices whose consequences are hard to anticipate, especially if an LLM is the author. We present a method for systematically comparing different formalizations of the same legal provision by their inferences on individual cases. Given multiple formalizations of a provision, we match them at the node level, derive a shared interface for each pair from the matching, and use a SAT solver to enumerate the edge cases on which any two formalizations disagree. Selected edge cases are then verbalized into concrete factual scenarios that a legal expert can examine and act on. We apply our method to formalizations of ten EU provisions generated by nine frontier LLMs. We find that behavioral divergence between formalizations is essentially uncorrelated with their structural agreement and that the verbalized cases reveal qualitatively distinct types of disagreement, including divergences that mirror genuine controversies in the legal commentary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。