arXiv:2608.10268cs.LGcs.AI2026-08

首个基于国际人权法的LLM评估基准,可测模型法律推理能力。

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

论文配图:Toward Human Rights Benchmarking for LLMs: A Pilot Methodology
图 1 · 摘自论文原文
  • 用IRAP框架重构法律推理,适配人权法独特逻辑
  • 模型在任务中准确率0.339至0.577,最低仅0.025
  • 适合关注AI伦理与法律合规的研究者使用

大型语言模型日益介入人权实现的法律判断,但尚无评估其人权法推理能力的基准。为此,我们提出构建首个专家验证、情景驱动的人权基准HumRightsBench。通过改进IRAC法律推理框架,以“提出补救措施”替代“法律结论”,形成IRAP结构化评估方法,并设计一系列真实世界人权情景,由全球人权律师与专业人士标注。结果显示,模型在不同任务上的准确率差异显著(整体表现0.339–0.577,任务范围0.025–0.774),表明该基准具备推动人工智能评估科学发展的潜力,恰逢其时。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench: the first expert-validated, scenario-based benchmark for evaluating reasoning grounded in the obligation structure of international human rights law. We adapt the IRAC framework for legal reasoning to better suit the unique reasoning patterns of human rights work (substituting P, "proposing remedies," for C, "legal conclusion," yielding IRAP) to structure our evaluation heuristics. We also produce a pilot series of authentic scenarios designed to implicate the many dimensions of real-world human rights issues and annotated by human rights lawyers and professionals across the world. Ultimately, we find that model accuracy scores range considerably across legal reasoning tasks (overall model performance ranges from 0.339 to 0.577, task min-max ranges from 0.025 to 0.774), which strongly implies that HumRightsBench is a capable instrument for advancing this emerging subfield of AI evaluations science at a critical moment in its evolution.

大模型评估人权法法律推理AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。