arXiv:2606.18021cs.AIcs.CL2026-06

法律AI幻觉问题严重,该研究提出可定位错误类型与方向的审计框架。

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

论文配图:LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI
图 1 · 摘自论文原文
  • 按法律类别分类幻觉,识别数值、时间、义务等四类错误模式
  • 发现同一模型在不同类别间幻觉率差距达38-40个百分点,平均值掩盖真实风险
  • 用诊断结果校准多智能体辩论,提升检测精度并降低虚构内容45%

部署于法律流程中的AI系统幻觉率高达约52%,但这一平均值掩盖了错误集中于特定类型及方向的问题,使合规人员难以获得可操作的风险信号。本文提出LegalHalluLens审计框架,包含三部分:基于CUAD数据集(Hendrycks et al., 2021)对四类法律驱动型主张(数值、时间、义务/权利、事实)的类型化幻觉画像;将遗漏与虚构偏差压缩为单一可比的部署指标——风险方向指数(RDI);以及针对误差幅度和方向进行校准的类型化辩论管道。在510份合同、249,252个条款级实例上,我们发现同一模型在义务/数值与时间类主张间的幻觉率差距约为38–40个百分点,而两个幻觉率均为52%的系统可能具有相反的RDI。辩论管道使虚假检测减少45%,各类型收益与诊断结果一致,且仅用40亿活跃参数的较小骨干模型即达到商用API水平。类型化画像与RDI揭示了聚合指标所隐藏的失效模式,并进一步证明这些诊断可作为多智能体辩论的校准输入,其中针对特定失效模式设计的质疑者与非对称门控机制优于通用调参。该框架支持方向感知的采购、问责与智能体设计,助力法律AI在真实场景中可信部署。

原文摘要 · Abstract (English)

AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they run, leaving compliance officers without an actionable signal for trustworthy deployment. We present LegalHalluLens, an auditing framework with three components: typed hallucination profiles across four legally-motivated claim categories (numeric, temporal, obligation/entitlement, factual) over CUAD (Hendrycks et al., 2021); a Risk Direction Index (RDI) that reduces omission-versus-invention bias to a single deployment-comparable scalar; and a typed debate pipeline calibrated to both magnitudes and directions. Across 510 contracts and 249,252 clause-level instances we measure a within-model gap of approximately 38-40 pp between obligation/numeric and temporal claims that aggregate reporting hides, and show that two systems with matched 52% rates can carry opposite RDIs. The debate pipeline reduces fabricated detections by 45% with per-category gains tracking the diagnosis, matching commercial APIs with a substantially smaller backbone (4B active parameters). Typed profiles and RDI surface failure modes that aggregate metrics hide; we further show these diagnostics serve as calibration inputs for multi-agent debate pipelines, where Skeptic challenges and asymmetric gates targeted at measured failure modes outperform generically-tuned debate. The framework supports direction-aware procurement, accountability, and agent design for legal AI deployed in the wild.

法律AI幻觉审计多智能体风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。