构建法律文档矛盾检测基准,提升大模型法律推理可靠性
LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents
- 多智能体框架生成带六类结构化矛盾的合成法律文本
- 结合自动挖掘与人工验证,确保矛盾真实可信
- 适用于法律RAG系统性能评估,推动可解释司法AI发展
检索增强生成(RAG)将大语言模型与外部信息源结合,但检索到的证据中未解决的矛盾常导致幻觉和法律上不成立的输出。现有矛盾检测基准缺乏领域真实性,仅覆盖有限冲突类型,且极少超出单句对,难以满足法律应用需求。因此,可控生成含矛盾的文档至关重要:它可系统性测试模型性能,覆盖多样冲突类别,并为矛盾检测与解决提供可靠评估基础。本文提出一种面向法律领域的多智能体矛盾感知基准框架,可生成具有法律风格的合成文档,注入六种结构化矛盾类型,并建模自洽性及成对不一致。通过自动化矛盾挖掘结合人工验证,保障内容合理性与真实性。该基准是法律RAG管道中少数结构化矛盾评估资源之一,有助于实现更一致、可解释、值得信赖的系统。
原文摘要 · Abstract (English)
Retrieval-Augmented Generation (RAG) integrates large language models (LLMs) with external sources, but unresolved contradictions in retrieved evidence often lead to hallucinations and legally unsound outputs. Benchmarks currently used for contradiction detection lack domain realism, cover only limited conflict types, and rarely extend beyond single-sentence pairs, making them unsuitable for legal applications. Controlled generation of documents with embedded contradictions is therefore essential: it enables systematic stress-testing of models, ensures coverage of diverse conflict categories, and provides a reliable basis for evaluating contradiction detection and resolution. We present a multi-agent contradiction-aware benchmark framework for the legal domain that generates synthetic legal-style documents, injects six structured contradiction types, and models both self- and pairwise inconsistencies. Automated contradiction mining is combined with human-in-the-loop validation to guarantee plausibility and fidelity. This benchmark offers one of the first structured resources for contradiction-aware evaluation in legal RAG pipelines, supporting more consistent, interpretable, and trustworthy systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。