评测大模型在多轮法律咨询中的交互能力,发现最佳模型仅达0.562得分。
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation

- 构建四类真实客户行为的多轮对话评测集
- 26个模型中最佳表现仅0.562,差距显著
- 揭示模型在需引导时反而表现更差的悖论
律师-客户咨询是法律服务的关键起点。有效法律协助依赖于从客户处获取充分且真实的信息,以制定最优保护策略。这要求大语言模型不仅具备扎实的法律推理能力,还需通过多轮互动策略性地提取关键事实,并有效应对不同性格的客户。然而现有法律评测基准忽视了这种交互能力。为此,我们提出DLawBench,一个面向真实法律咨询的诊断性评测基准。基于真实客户行为,将律师-客户互动分为四类:合作型、依赖型、退缩型和对抗型。采用真实案例对话,评估大模型在现实条件下的法律咨询能力。DLawBench包含461个中、美法律案例,5,532对事实条目,3,411个提问规范,3,348个问题-解决规范,评测了26个代表性大模型。系统实验显示巨大提升空间:表现最好的GPT-5.5在咨询基础法律推理上仅得0.562。更重要的是,该基准暴露了模型在法律咨询中的奉承现象及悖论:当客户最需要引导时,模型表现反而最差。
原文摘要 · Abstract (English)
Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order to devise strategies that best protect their interests. This task requires Large Language Models (LLMs) not only to perform robust legal reasoning, but also to strategically elicit material facts through multi-turn interactions and effectively guide clients with diverse personalities. Yet existing legal benchmarks overlook this interactive capability. To fill this gap, we introduce DLawBench, a diagnostic benchmark for real-world legal consultation. Drawing on realistic client behavior, we characterize lawyer-client interactions into four types: Cooperative, Dependent, Withdrawn, and Adversarial. Using dialogues grounded in real cases, DLawBench evaluates whether LLMs can effectively conduct legal consultation under realistic conditions. DLawBench comprises 461 cases from Chinese and U.S. law, 5,532 paired fact entries, 3,411 inquiry rubrics, and 3,348 issue-resolution rubrics, and evaluates 26 representative LLMs. Systematic experiments show substantial headroom: the best-performing model, GPT-5.5, achieves only 0.562 on consultation-grounded legal reasoning. More importantly, DLawBench exposes both sycophancy in legal consultation and a paradox: models perform worse when clients need guidance most.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。