arXiv:2605.05532cs.CLcs.CY2026-05被引 1

小模型在合同结构化提取上表现超越大模型,成本降九成

A Few Good Clauses: Comparing LLMs vs Domain-Trained Small Language Models on Structured Contract Extraction

论文配图:A Few Good Clauses: Comparing LLMs vs Domain-Trained Small Language Models on Structured Contract Extraction
图 1 · 摘自论文原文
  • 用领域微调的小模型自研部署,替代大模型推理
  • 宏平均F1达0.812,推理成本降低78%至97%
  • 误报少,适合法律等高风险场景应用

本文评估领域训练的小语言模型(SLM)是否能在结构化合同提取任务上超越前沿大语言模型,且成本大幅降低。实验对比了自托管法律领域专家混合模型Olava Extract与五款前沿模型。Olava Extract在各项指标中表现最优,宏平均F1为0.812,微平均F1为0.842,相较测试的前沿模型推理成本降低78%至97%。其精确率最高,生成更少幻觉和无支持的提取结果,这对法律流程中避免操作风险和后续审查负担至关重要。研究结果表明,高性能、可媲美人工的法律AI不再依赖最大规模的外部托管模型。更广泛而言,挑战了企业级高价值AI必须依赖超大规模模型、巨额基础设施投入及中心化服务商的固有假设。

原文摘要 · Abstract (English)

This paper evaluates whether a domain trained Small Language Model (SLM) can outperform frontier Large Language Models on structured contract extraction at radically lower cost. We test Olava Extract, a self hosted legal domain Mixture of Experts model, against five frontier models. Olava Extract achieved the strongest aggregate performance in the study, with a macro F1 of 0.812 and a micro F1 of 0.842, while reducing inference cost by 78% to 97% compared with the frontier models tested. It also achieved the highest precision scores, producing fewer hallucinated and unsupported extractions, an important distinction in legal workflows where hallucinations create operational risk and downstream review burden. The findings shows that high performing, human comparable legal AI no longer requires the largest externally hosted models. More broadly, they challenge the assumption that commercially valuable enterprise AI capability must remain tied to ever larger models, massive infrastructure expenditure, and centrally hosted providers.

法律AI小模型合同抽取低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。