用大模型指导小模型,高效提升法律长文摘要的论点覆盖度
Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation

- 通过序列级蒸馏,让小模型从大模型获取高质量监督信号
- 仅需约10个样本即可显著提升论点覆盖率,数据效率高
- 适合需要低成本、高精度法律文本摘要的场景
我们证明,从具备长上下文能力的教师模型进行序列级蒸馏,是一种简单、无需标注、数据高效的策略,可有效提升小规模语言模型在长篇法律意见摘要中的论点重要性覆盖度。在不同学生模型规模下,蒸馏方法均优于基于专家撰写摘要的微调。进一步实验表明,绝大多数性能提升可在仅使用约10个训练摘要的情况下实现,凸显了教师生成监督信号的强数据效率。最后发现,仅摘要蒸馏已足够带来显著改进:推理链蒸馏虽表现相当,但与摘要监督结合后收益微弱。
原文摘要 · Abstract (English)
We show that sequence-level distillation from a capable long-context teacher model is a simple, annotation-free, and data-efficient strategy for improving argument saliency coverage in long legal opinion summarization, where small LLMs often struggle to retain the most salient argumentative content. Across student model sizes, distillation consistently surpasses tuning on expert-written summaries in our legal-opinion setting. We further demonstrate that most gains are achieved with as few as ~10 training summaries, highlighting the strong data efficiency of teacher-generated supervision. Finally, we find that summary distillation is sufficient for improvements: reasoning-chain distillation remains competitive with summary-only distillation, but provides marginal benefit when combined with summary supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。