arXiv:2607.11400cs.CL2026-07

用多模型集成与提示工程提升法律信息处理精度,多项任务夺冠。

Cross-Architecture LLM Ensembles, Feature-Based Reranking and Retrieval-Augmented Prompting for Legal Information Processing

  • 跨架构模型集成+特征重排序+检索增强提示,适配不同法律任务
  • 法规蕴含任务达96.3%准确率,超越33个参赛团队
  • 通过优化提示词,推理性能提升超60%,适合法律AI研究者参考

法律信息处理涵盖检索、蕴含判断和判决预测,需文本匹配、推理与弱监督下的鲁棒泛化。本研究为团队DU在COLIEE 2026全部五项任务中的参与成果,采用开源大模型完成法律案例检索、案例蕴含、法规检索与蕴含、以及法律判决预测。任务3与4的模型均在2025年7月15日前发布。任务4(法规蕴含)中,九个来自三个模型家族的跨架构集成模型达到96.3%准确率,位列33个提交方案中的第一名。试点任务(侵权预测与理由提取)中,结合五个主张级模型并利用主张预测特征优化判决结果,实现73.1%的真阳性准确率与68.2%的召回精确率,非官方提交中优于所有官方结果,在真阳性上领先,召回精确率并列最高。任务2(法律案例蕴含)中,仅将提示从单选改为多选,后验评估显示F1从0.343提升至0.555,超过最佳官方提交(F1=0.490)。任务3(法规检索与蕴含)中,以Qwen3-235B替换原蕴含模型,并使用结构化法律推理提示,准确率从79.3%提升至91.5%。任务1(法律案例检索)中,结合词汇与语义检索,融合34项结构、引用权威性与时间特征的排序系统,取得F1=0.314(54个提交中排名第11)。整体表明,不同任务受益于各异归纳偏置,跨架构集成、基于特征的重排序与检索增强提示在不同场景中表现最优。

原文摘要 · Abstract (English)

Legal information processing spans retrieval, entailment and judgment prediction problems, requiring text matching, reasoning and robust generalisation with limited supervision. We report Team DU's participation in all five tasks of COLIEE 2026, using open-weight systems for legal case retrieval, case entailment, statute retrieval and entailment, and legal judgment prediction. For Tasks 3 and 4, all models predate the 15 July 2025 cutoff required by the rules. For Task 4 (statute entailment), a cross-architecture ensemble of nine models from three families achieves 96.3% accuracy, placing first among 33 submissions from 11 teams. For the Pilot Task (tort prediction and rationale extraction), a multi-view system combining five claim-level models and refining the verdict using features derived from the claim predictions achieves 73.1% TP accuracy and 68.2% RE F1 as an unofficial submission, scoring above all official entries on TP and matching the highest on RE. For Task 2 (legal case entailment), changing only the prompt from single- to multi-selection raises F1 from 0.343 to 0.555 in post-competition evaluation on released gold labels, exceeding the best official submission (F1 = 0.490). For Task 3 (statute retrieval and entailment), replacing the entailment model with Qwen3-235B and a structured legal reasoning prompt raises accuracy from 79.3% to 91.5% in post-competition analysis. For Task 1 (legal case retrieval), a learning-to-rank system combining lexical and semantic retrieval with structural, citation authority, and temporal features (34 in total) achieves F1 = 0.314 (rank 11 of 54 submissions from 22 teams). Overall, legal information processing benefits from different inductive biases across tasks, with cross-architecture ensembling, feature-based reranking and retrieval-augmented prompting each proving most effective in different settings.

法律AI模型集成提示工程信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。