用动态生成任务诊断法律推理模型短板,精准定位能力缺陷。
OpenExempt: A Diagnostic Benchmark for Legal Reasoning and a Framework for Creating Custom Benchmarks on Demand
- 用符号化法规生成可定制的法律推理题,支持按需构建测试集。
- 9765个样本分九类,发现长推理链和干扰句会引发性能骤降。
- 适合研究法律推理、模型可解释性或想自建评测基准的团队。
推理评测对语言模型发展至关重要,但静态问答仅提供性能快照,难以揭示复杂行为。在规则密集型领域如法律,现有评测成本高且难定位具体失败模式。为此,我们提出OpenExempt框架与基准:利用专家设计的美国破产法条文符号表示,动态生成大量自然语言推理任务及其机器可计算解。用户可精细控制任务复杂度与范围,独立探测特定推理能力。基于此,构建了包含9765个样本的OpenExempt基准,覆盖九个评估套件,系统测试13种不同语言模型,发现模型在长推理路径和干扰语句存在时出现明显性能断崖。框架与基准已公开,助力下一代推理系统研究。
原文摘要 · Abstract (English)
Reasoning benchmarks have played a crucial role in the progress of language models. Yet rigorous evaluation remains a significant challenge as static question-answer pairs provide only a snapshot of performance, compressing complex behavior into a single accuracy metric. This limitation is especially true in complex, rule-bound domains such as law, where existing benchmarks are costly to build and ill suited for isolating specific failure modes. To address this, we introduce OpenExempt, a framework and benchmark for diagnostic evaluation of legal reasoning. The OpenExempt Framework uses expert-crafted symbolic representations of U.S. Bankruptcy Code statutes to dynamically generate a large space of natural language reasoning tasks and their machine-computable solutions on demand. This gives users fine-grained control over task complexity and scope, allowing individual reasoning skills to be probed in isolation. Using this system, we construct the OpenExempt Benchmark, a diagnostic benchmark for legal reasoning with 9,765 samples across nine evaluation suites designed to carefully probe model capabilities. Experiments on 13 diverse language models reveal sharp performance cliffs that emerge only under longer reasoning paths and in the presence of obfuscating statements. We release the framework and benchmark publicly to support research aimed at understanding and improving the next generation of reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。