用教师模型生成高质量法律推理数据,提升小模型能力
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models

- 通过细粒度提示从大模型提取推理轨迹
- 自反思验证筛选有效数据,使小模型性能显著提升
- 无需专家标注,适合资源有限的法律AI落地
小型语言模型(SLMs)因高效与低成本,适合实际部署。但其容量有限,在需要连贯法条解读和逻辑一致推演的高风险法律任务中表现不佳。此外,训练 SLM 所需的高质量、简洁推理轨迹,人工收集成本过高,常规拒采方法也难以获取足够细粒度的标注,仅能提供最终判决结果。为此,我们提出 {LegalDrill}——一种诊断驱动的数据合成框架:从高性能教师模型中通过细粒度提示提取并迭代优化推理轨迹,再通过自反思验证机制,动态筛选最有效的数据用于小模型学生训练。这些数据支持监督微调与直接偏好优化。在多个法律基准上的实验表明,{LegalDrill} 显著增强了代表性 SLM 的法律推理能力,且无需依赖稀缺专家标注,为实用化法律推理系统提供了可扩展路径。
原文摘要 · Abstract (English)
Small language models (SLMs) are promising for real-world deployment due to their efficiency and low operational cost. However, their limited capacity struggles with high-stakes legal reasoning tasks that require coherent statute interpretation and logically consistent deduction. Furthermore, training SLMs for such tasks demands high-quality, concise reasoning trajectories, which are prohibitively expensive to manually collect and difficult to curate via standard rejection sampling, lacking granularity beyond final verdicts. To address these challenges, we propose {LegalDrill}, a diagnosis-driven synthesis framework that extracts and iteratively refines reasoning trajectories from a capable teacher via fine-grained prompting, then a self-reflective verification is employed to adaptively select the most effective data for the SLM student. The resulting data empower SLM training through supervised fine-tuning and direct preference optimization. Extensive experiments on several legal benchmarks demonstrate that {LegalDrill} significantly bolsters the legal reasoning capabilities of representative SLMs while bypassing the need for scarce expert annotations, paving a scalable path toward practical legal reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。