arXiv:2608.24936cs.LGcs.AI2026-08

0.6B参数小模型,专为法律检索优化,性能媲美大模型。

GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval

  • 两阶段训练:先蒸馏大模型知识,再用硬负样本微调
  • 在MLEB上达75.11%,MTEB(Law)上达64.38%
  • 支持多种量化格式,适合低资源部署场景

我们提出GreenLeaf Law Embed Tiny,一个仅0.6B参数的法律领域检索嵌入模型。该模型在Massive Legal Embedding Benchmark(MLEB)上取得75.11%的准确率,在MTEB(Law, v1)上达到64.38%,在参数量低于1B的模型中表现优异。方法包括:两阶段训练流程——先从大模型蒸馏知识到紧凑学生架构,再通过硬负样本挖掘进行领域特定微调;构建包含340万条查询-段落对的数据集,涵盖15万条跨司法管辖区的人工标注样本;采用高效推理架构,支持BF16、INT8和二值化等多种量化级别,可在资源受限环境中部署。我们还对训练方法、结构设计及法律检索任务进行了全面评估。结果表明,高质量数据与领域定制训练可显著提升专用场景下的性能。

原文摘要 · Abstract (English)

We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models under 1B parameters. Our approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies domain-specific fine-tuning with hard negative mining; a carefully curated dataset of 3.4 million query-passage pairs, including 150,000 human-curated samples across diverse legal jurisdictions; and an efficient inference architecture supporting multiple quantization levels (BF16, INT8, binary) enabling deployment in resource-constrained environments. We provide detailed analysis of our training methodology, architectural choices, and comprehensive evaluation across legal retrieval tasks. Our results demonstrate that domain-specific training with high-quality data can improve performance for specialized domain applications

法律检索小模型嵌入模型量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。