arXiv:2504.02881cs.CL2025-04

LLM在法律账单审核中全面超越人类,效率提升百倍且成本下降99.97%

Better Bill GPT: Comparing Large Language Models against Legal Invoice Reviewers

  • 用真实法律专家标注的账单数据,对比LLM与人类评审员表现
  • LLM准确率达92%,处理速度仅3.6秒/张,人力需3至5分钟
  • 适合法律科技、企业法务及成本控制团队参考应用

法律账单审核耗时耗力且标准不一,传统上由初级律师、资深律师或法务专员逐条核查。本研究首次对大型语言模型(LLMs)与人类评审员进行实证比较,评估其准确性、速度和成本效益。基于专家法律人士设定的基准数据集,结果显示:在账单审批决策中,LLM最高准确率达92%,远超资深律师72%的上限;在明细项分类上,顶级模型F-score达81%,人类最佳组仅43%;处理速度方面,律师平均需194至316秒/份,而LLM最快仅需3.6秒;成本方面,人工平均每单4.27美元,AI降至不足1美分,降幅达99.97%。该研究揭示了人工智能在法律支出管理中的变革性作用,表明大模型驱动的法律审核时代已然到来。未来挑战已非能否替代人类,而是如何战略整合自动化与人工判断。

原文摘要 · Abstract (English)

Legal invoice review is a costly, inconsistent, and time-consuming process, traditionally performed by Legal Operations, Lawyers or Billing Specialists who scrutinise billing compliance line by line. This study presents the first empirical comparison of Large Language Models (LLMs) against human invoice reviewers - Early-Career Lawyers, Experienced Lawyers, and Legal Operations Professionals-assessing their accuracy, speed, and cost-effectiveness. Benchmarking state-of-the-art LLMs against a ground truth set by expert legal professionals, our empirically substantiated findings reveal that LLMs decisively outperform humans across every metric. In invoice approval decisions, LLMs achieve up to 92% accuracy, surpassing the 72% ceiling set by experienced lawyers. On a granular level, LLMs dominate line-item classification, with top models reaching F-scores of 81%, compared to just 43% for the best-performing human group. Speed comparisons are even more striking - while lawyers take 194 to 316 seconds per invoice, LLMs are capable of completing reviews in as fast as 3.6 seconds. And cost? AI slashes review expenses by 99.97%, reducing invoice processing costs from an average of $4.27 per invoice for human invoice reviewers to mere cents. These results highlight the evolving role of AI in legal spend management. As law firms and corporate legal departments struggle with inefficiencies, this study signals a seismic shift: The era of LLM-powered legal spend management is not on the horizon, it has arrived. The challenge ahead is not whether AI can perform as well as human reviewers, but how legal teams will strategically incorporate it, balancing automation with human discretion.

法律AI大模型应用成本优化智能审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。