arXiv:2506.04636cs.AIcs.CL2025-06被引 2

测试大模型在公司治理规则下的推理能力,发现现有模型表现有限。

CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models

  • 构建首个基于真实公司章程的法律推理基准,判断高管行为是否合规。
  • GPT-4o准确率75.2%,顶级推理代理最高达78.1%,仍存明显差距。
  • 适合研究法律AI、模型推理能力评估的学者和工程师参考。

法律是自然语言处理的重要应用领域,而推理(尤其是对判例的关联能力)是法律实践的核心。尽管已有多个法律数据集,但尚无专攻推理任务的基准。本文提出针对公司治理领域的推理评测基准CHANCERY,模拟真实世界公司治理法规,评估模型判断高管/董事会/股东提案是否符合公司章程的能力。该基准基于24条公认公司治理原则及79份真实行业章程(从10,000份中筛选),要求模型进行二分类判断。对当前最优推理模型的测试表明其难度极高:Claude 3.7 Sonnet与GPT-4o准确率分别为64.5%和75.2%。采用ReAct与CodeAct框架的推理代理分别取得76.1%与78.1%的成绩,进一步揭示了高分需具备先进法律推理能力。此外,分析显示当前模型在复杂逻辑与多层条款关联上仍存在显著短板。

原文摘要 · Abstract (English)

Law has long been a domain that has been popular in natural language processing (NLP) applications. Reasoning (ratiocination and the ability to make connections to precedent) is a core part of the practice of the law in the real world. Nevertheless, while multiple legal datasets exist, none have thus far focused specifically on reasoning tasks. We focus on a specific aspect of the legal landscape by introducing a corporate governance reasoning benchmark (CHANCERY) to test a model's ability to reason about whether executive/board/shareholder's proposed actions are consistent with corporate governance charters. This benchmark introduces a first-of-its-kind corporate governance reasoning test for language models - modeled after real world corporate governance law. The benchmark consists of a corporate charter (a set of governing covenants) and a proposal for executive action. The model's task is one of binary classification: reason about whether the action is consistent with the rules contained within the charter. We create the benchmark following established principles of corporate governance - 24 concrete corporate governance principles established in and 79 real life corporate charters selected to represent diverse industries from a total dataset of 10k real life corporate charters. Evaluations on state-of-the-art (SOTA) reasoning models confirm the difficulty of the benchmark, with models such as Claude 3.7 Sonnet and GPT-4o achieving 64.5% and 75.2% accuracy respectively. Reasoning agents exhibit superior performance, with agents based on the ReAct and CodeAct frameworks scoring 76.1% and 78.1% respectively, further confirming the advanced legal reasoning capabilities required to score highly on the benchmark. We also conduct an analysis of the types of questions which current reasoning models struggle on, revealing insights into the legal reasoning capabilities of SOTA models.

法律AI推理评测公司治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。