arXiv:2607.17745cs.AI2026-07

构建环境执法大模型评测基准,揭示其推理短板

WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement

论文配图:WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement
图 1 · 摘自论文原文
  • 基于真实执法案例构建2521个任务,覆盖全流程环节
  • 模型在规则任务表现好,但证据链构建与多源整合能力弱
  • 适合研究法律AI、智能执法系统及模型评估的学者使用

大型语言模型(LLMs)在环境执法中日益受到关注,但其生成可追溯执法决策的能力仍不明确。我们提出WuYu-EnvLE-Bench,一个基于真实执法案例、法规标准与专家评审构建的基准,包含2,521个评测实例、14项任务和12个污染介质子领域,覆盖事前、事中、事后全流程。采用绝对环境执法评分(AES)与智能执法指数(IEI),评估开源与闭源模型在能力、响应质量与资源效率上的表现。结果表明,模型在规则限定任务中表现良好,但在证据链构建、矛盾检测、多源信息整合及程序性判断上仍不可靠;模型规模扩大也呈现边际效益递减:中等规模模型在结构化任务中已接近领先模型,更大模型未能有效突破证据推理瓶颈。该基准凸显了环境执法推理需具备证据锚定、规则敏感与任务自适应特性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.

大模型评测环境执法法律AI证据推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。