对比大模型与符号系统,发现法律推理需避免数据污染影响
Reasoners or Translators? Contamination-aware Evaluation and Neuro-Symbolic Robustness in Tax Law

- 用污染检测法验证大模型真实推理能力
- 混合系统在未见案例上准确率提升37%
- 适合法律AI可靠性研究者参考
大型语言模型(LLMs)在自动化法律推理方面取得进展,但其性能是否反映真实法律推理能力,仍受数据污染影响。本文开展全面实证研究,提出污染检测协议以严格评估模型可靠性。结果表明,性能可能因数据污染被夸大。在此基础上,系统比较了单一LLM与将法规文本翻译为形式化表示并交由符号求解器推理的混合系统。构建新型测试集,通过案例与规则变异探测泛化能力。研究发现法律推理具有内在组合性,神经符号框架在未观测情境中表现更稳健,具备更高可靠性与泛化能力。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have significantly enhanced automated legal reasoning. Yet, it remains unclear whether their performance reflects genuine legal reasoning ability or artifacts of data contamination. We present a comprehensive empirical study of tax law reasoning approaches and implement a contamination detection protocol to rigorously assess LLM reliability. We show that performance can be inflated by contamination. Building on this analysis, we conduct a systematic evaluation, comparing monolithic LLMs with hybrid systems that translate statutory text into formal representations and delegate inference to symbolic solvers. We build a novel test suite designed to probe generalization to unseen documents via case and rule variations. Our findings indicate that legal reasoning is inherently compositional and that neuro-symbolic frameworks offer a more reliable and robust foundation for legal AI, as well as improved generalization to unobserved situations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。