让规则自动进化,提升法律案例检索准确率
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval

- 用大模型自动生成并优化查询重写规则
- 在中文法律数据集上超越人工规则和贪心策略
- 适合法律AI、信息检索研究者参考
法律案例检索因法律语言复杂性和查询与案例间精确词汇对齐需求而极具挑战。尽管密集检索模型取得显著进展,实证研究仍表明BM25在该领域保持强大基线性能。这促使我们提出一种无需参数训练的自演化规则驱动查询重写框架。该框架赋予基于LLM的智能体自动评估环境,使其能迭代生成重写规则、规划规则组合的验证实验,并根据历史反馈淘汰无效规则。我们在中文法律案例检索基准LeCaRD-v2上评估该方法。实验结果表明,所提框架优于非演化基线,包括人工设计规则和贪心规则选择,尤其在使用高容量核心LLM时表现更优。我们还进行了详细分析,揭示了大模型利用过往实验结果及内在规则剔除知识在自演化中优化规则集的关键作用。
原文摘要 · Abstract (English)
Legal case retrieval remains challenging due to the complexity of legal language and the need for precise lexical alignment between queries and relevant cases. Although dense retrieval models have achieved notable progress, empirical studies show that BM25 continues to serve as a strong baseline in this domain. It motivates us to propose a self-evolving framework for rule-driven query rewriting that enhances BM25 without any parameter training. The framework equips an LLM-based agent with an automatic evaluation environment, enabling it to iteratively create rewriting rules, plan validation experiments over rule combinations, and eliminate ineffective rules based on historical feedbacks. We evaluate our method on the Chinese legal case retrieval benchmark LeCaRD-v2. Experimental results demonstrate that the proposed framework outperforms non-evolutionary baselines, including human-designed rules and greedy rule selection, particularly when powered by a highcapacity core LLM. We also conduct detailed analyses to investigate the mechanisms underlying self-evolution. Our findings reveal that LLM's capabilities to leverage previous experimental results and its intrinsic knowledge of rule elimination play critical roles in refining the rule set via self-evolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。