构建真实法律咨询对话数据集,评估大模型的交互与专业能力。
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
- 从短视频平台采集真实多轮法律咨询对话,覆盖3696条对话
- 专家标注确保专业性,模型在澄清能力上仅达39.8%召回率
- 适合研究法律对话系统、大模型司法应用的学者与开发者
法律咨询对保障个人权利和实现司法公正至关重要,但因专业人才短缺,普遍存在成本高、可及性差的问题。尽管大语言模型(LLMs)为低成本、可扩展的法律援助带来希望,现有系统仍难以应对真实咨询中复杂的交互与知识需求。为此,我们推出LeCoDe——一个基于真实场景的多轮法律咨询对话基准数据集,包含3,696条对话、共110,008个对话轮次。数据通过短视视频平台直播采集,保证真实性;并由法律专家严格标注,融入专业判断。我们提出一个综合评估框架,从(1)澄清能力与(2)专业建议质量两个维度,采用12项指标全面评测模型表现。在多种通用与领域专用大模型上的实验显示,即使先进模型如GPT-4,在澄清能力上也仅达39.8%召回率,整体建议质量得分59%,凸显该任务复杂性。基于此,我们进一步探索提升策略。本基准推动法律对话系统研究,更贴近真实人机法律互动场景。
原文摘要 · Abstract (English)
Legal consultation is essential for safeguarding individual rights and ensuring access to justice, yet remains costly and inaccessible to many individuals due to the shortage of professionals. While recent advances in Large Language Models (LLMs) offer a promising path toward scalable, low-cost legal assistance, current systems fall short in handling the interactive and knowledge-intensive nature of real-world consultations. To address these challenges, we introduce LeCoDe, a real-world multi-turn benchmark dataset comprising 3,696 legal consultation dialogues with 110,008 dialogue turns, designed to evaluate and improve LLMs' legal consultation capability. With LeCoDe, we innovatively collect live-streamed consultations from short-video platforms, providing authentic multi-turn legal consultation dialogues. The rigorous annotation by legal experts further enhances the dataset with professional insights and expertise. Furthermore, we propose a comprehensive evaluation framework that assesses LLMs' consultation capabilities in terms of (1) clarification capability and (2) professional advice quality. This unified framework incorporates 12 metrics across two dimensions. Through extensive experiments on various general and domain-specific LLMs, our results reveal significant challenges in this task, with even state-of-the-art models like GPT-4 achieving only 39.8% recall for clarification and 59% overall score for advice quality, highlighting the complexity of professional consultation scenarios. Based on these findings, we further explore several strategies to enhance LLMs' legal consultation abilities. Our benchmark contributes to advancing research in legal domain dialogue systems, particularly in simulating more real-world user-expert interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。