arXiv:2608.17168cs.CLcs.AI2026-08

测试大模型在人权法院案件中的法律推理能力,发现其分析表面完整但实质浅显。

Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

论文配图:Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases
图 1 · 摘自论文原文
  • 用欧洲人权法院案例测试大模型推理,对比不同提示策略效果
  • 模型生成结构完整但内容浅薄的法律分析,准确率未因提示优化提升
  • 自评模型结果可靠但不准确,不能替代人工评估,警惕误用准确率衡量推理

推理已成为当代大语言模型的标准技术与特征,但在法律导向任务(如案件预测)中的应用与质量仍缺乏深入研究。本文以欧洲人权法院(ECtHR)案例为测试基准,评估OpenAI GPT 5.4在法律案件预测中的推理能力,探索多种提示策略对法律有意义推理的影响。通过人类与大模型双重评估,发现该模型在法律推理上表现远未达理想水平:虽生成结构完整的分析,但实质内容浅显;采用专家定制提示虽提升分析全面性,但未带来预测准确率提升。此外,以大模型作为裁判者的评估具内部一致性,但与专业标注者相关性弱,即结果可靠但无效。研究呼吁避免仅依赖自动化评估,切勿将任务准确率视为推理质量的合理代理。

原文摘要 · Abstract (English)

Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substitute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality.

法律推理大模型评估ECtHR提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。