arXiv:2606.19812cs.AIcs.LG2026-06

用人工介入机制防止AI在法律文档审查中因错误累积导致失密风险。

Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery

  • 构建法律信息检索中AI代理失效的分阶段分类体系。
  • 通过四层验证架构将误判风险降低61%,仅需少于1/4文档人工复核。
  • 适合法律科技公司和合规团队参考,防范自动化审查的致命缺陷。

自主大语言模型(LLM)代理正被广泛应用于电子发现(e-discovery),但多步推理链中的错误叠加可能构成法律过失。与单次检索不同,基于敏感文档库的代理工作流存在一种称为“轨迹坍缩”的失效模式:早期误判会无声传播,使整个保密审查失效。本文提出三项贡献:首先,构建了法律信息检索中代理失效的结构化分类体系,按功能阶段划分;其次,设计了一套四层验证架构,涵盖规划、推理、执行与不确定性量化,可在错误累积前及时拦截;第三,基于合成e-discovery语料的初步仿真研究显示,强制设置人工介入(HOTL)阈值可显著降低保密豁免风险。结果表明,经过校准的不确定性阈值能使保密豁免风险相较完全自治部署降低高达61%,同时仅需将不到四分之一的文档转交律师审查。

原文摘要 · Abstract (English)

Autonomous Large Language Model (LLM) agents are increasingly deployed in electronic discovery (e-discovery), where compounding errors across multi-step reasoning chains can constitute legal malpractice. Unlike single-turn retrieval, agentic workflows operating over privileged document corpora exhibit a class of failure we term "trajectory collapse": an early misclassification silently propagates, rendering an entire privilege review invalid. This paper makes three contributions. First, we propose a structured taxonomy of agentic failures in legal information retrieval, organized by functional stage. Second, we introduce a four-layer verification architecture -- spanning planning, reasoning, execution, and uncertainty quantification -- designed to intercept these failures before they compound. Third, we present a preliminary simulation study on a synthetic e-discovery corpus that demonstrates how mandatory Human-on-the-Loop (HOTL) escalation thresholds reduce privilege-waiver risk relative to fully autonomous baselines. Our results suggest that calibrated uncertainty thresholds can reduce privilege-waiver risk by up to 61% versus fully autonomous deployment, while routing fewer than one quarter of documents to attorney review.

法律AI人机协作大模型风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。