构建韩语法律问答基准,提升多跳推理能力
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
- 用大模型与专家结合生成法律场景问答对
- 提出参数化条款引导检索方法,准确率提升37.91点
- 专设法律真实性评估指标,适配法律领域研究者
大型语言模型(LLM)在通用领域表现优异,并逐步拓展至法律专业领域。现有评测基准难以评估开放性、基于法条的问答任务。为此,我们提出韩语法律可解释问答基准KoBLEX,包含226个基于场景的问答实例及其支持法条,通过大模型与人工专家混合流程构建。我们还提出参数化条款引导检索(ParSeR)方法,利用大模型生成的参数化法条引导法律相关且可靠的答句。ParSeR通过三阶段顺序检索机制,实现复杂法律问题的多跳推理。为进一步评估答案的法律一致性,我们设计了法律真实性评估(LF-Eval),该自动指标综合考虑问题、答案和支撑法条,与人工判断高度相关。实验表明,ParSeR在多个主流大模型上均优于强基线,相较于GPT-4o标准检索,F1提升37.91,LF-Eval提升30.81。消融分析证实其在不同推理深度下性能稳定有效。
原文摘要 · Abstract (English)
Large Language Models (LLM) have achieved remarkable performances in general domains and are now extending into the expert domain of law. Several benchmarks have been proposed to evaluate LLMs' legal capabilities. However, these benchmarks fail to evaluate open-ended and provision-grounded Question Answering (QA). To address this, we introduce a Korean Benchmark for Legal EXplainable QA (KoBLEX), designed to evaluate provision-grounded, multi-hop legal reasoning. KoBLEX includes 226 scenario-based QA instances and their supporting provisions, created using a hybrid LLM-human expert pipeline. We also propose a method called Parametric provision-guided Selection Retrieval (ParSeR), which uses LLM-generated parametric provisions to guide legally grounded and reliable answers. ParSeR facilitates multi-hop reasoning on complex legal questions by generating parametric provisions and employing a three-stage sequential retrieval process. Furthermore, to better evaluate the legal fidelity of the generated answers, we propose Legal Fidelity Evaluation (LF-Eval). LF-Eval is an automatic metric that jointly considers the question, answer, and supporting provisions and shows a high correlation with human judgments. Experimental results show that ParSeR consistently outperforms strong baselines, achieving the best results across multiple LLMs. Notably, compared to standard retrieval with GPT-4o, ParSeR achieves +37.91 higher F1 and +30.81 higher LF-Eval. Further analyses reveal that ParSeR efficiently delivers consistent performance across reasoning depths, with ablations confirming the effectiveness of ParSeR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。