L-MARS提升法律问答引用真实性,通过多智能体验证与检索修复。
L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit
- 采用多智能体协作搜索与判官式证据核查机制
- 引用准确率从0.13升至0.25,无引文率降至13%
- 适合法律AI研究者与注重可信生成的实践应用
大型语言模型在法律问答中的应用日益广泛,但现有评估多关注选择题准确率,忽视了答案所附引文是否真实存在且支持所述规则。本文提出L-MARS,一个开源多智能体法律问答系统,融合代理式搜索与判官驱动的证据审计。每个原子命题按六类标签分类,并在跨模型家族的判官评估下使用strict-ALCE打分。在分层的100题法考审计中,检索对准确率影响微弱,但多轮判官循环使严格引用F1从0.13(原始RAG)提升至0.25,无引文率由34%降至13%。我们引入Faith-Search作为草案后验证与修复步骤,将不可达引文率压至1%以下,但未进一步提升F1,故视为针对性可达性干预而非忠实度突破。50题LegalSearchQA案例研究证实:检索-生成流水线在外部审计下引用F1趋近0.75,而单智能体网络搜索基线骤降至0.22。
原文摘要 · Abstract (English)
Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy. This measure overlooks a common failure: whether the citation source attached to an answer exists and supports the rule the system attributes to it. We present L-MARS, an open multi-agent legal QA system with agentic search and judge-driven evidence checks, and audit it claim by claim against its cited source. Each atomic claim is labelled with a six-class taxonomy and scored with strict-ALCE under cross-provider judging, where the answerer and verifier come from different model families. On a stratified 100-question Bar Exam audit, retrieval barely moves accuracy, yet the multi-turn judge loop lifts strict citation F1 from 0.13 (naive RAG) to 0.25 and cuts the no-citation rate from 34% to 13%. We further introduce Faith-Search, a post-draft step that re-verifies and repairs unreachable citations; it drops the unreachable rate below 1% but does not improve F1 over the multi-turn loop, so we report it as a targeted reachability intervention rather than a faithfulness breakthrough. A 50-question LegalSearchQA case study confirms the picture: retrieve-then-draft pipelines saturate near 0.75 citation F1, while a single-agent web-search baseline collapses to 0.22 under external audit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。