小模型在法律问答中真能用好给的法条吗?实验证明它们常依赖检索而非真正理解。
Do Small Models Use the Law You Give Them? Measuring Context Use on a Bilingual Bangladesh Legal Benchmark
- 设计双语孟加拉法律数据集,区分模型、检索与评分的影响
- 移除法条后模型准确率下降8%-15%,但微调未提升使用法条的能力
- 发现模型看似进步实则可能来自评分方式偏差,适合法律AI评估研究者
微调可提升法律问答准确率,但未必改善模型对上下文法条的实际利用。我们在双语孟加拉法律问答任务中研究这一差异,观察到错误可能源于答案评分、检索或未能使用相关法条。构建了包含2,165个经审核的双语微调样本的层级保留法条语料库,以及150项提供法条的对照测试集。评估六种指令微调模型:Llama-3.2-1B、Llama-3.2-3B、Qwen3.5-0.8B、Qwen3.5-2B、Qwen3.5-4B 和 Gemma-4-E2B,每模型三组LoRA种子。通过约束选项字母评分、循环选项轮换和受控移除管辖条款的方法分离影响。在398个律师理事会输出上,精确行解析显示,Qwen3.5-2B seed-42适配器带来50.0%的准确率提升,而选项评分仅得3.0%。Gemma-4-E2B在两种评分下偏好不同系统。当确保管辖条款存在时,六种参考模型在四阶标准下准确率提升14.7%–19.3%;移除该条款后,模型准确率下降8.0%–15.3%,适配器下降13.8%–14.9%。但差分估算显示微调后模型对法条的依赖并未增加。结果表明,法律适应性声称需分离评分者、检索者与模型效应。代码与数据见https://anonymous.4open.science/r/bangladesh-legal-qa-11E3
原文摘要 · Abstract (English)
Fine-tuning can improve legal question-answering accuracy without improving how models use law supplied in context. We study this distinction in bilingual Bangladeshi legal QA, where observed errors can arise from answer scoring, retrieval, or failure to use relevant law. We construct a hierarchy-preserving statutory corpus, 2,165 reviewed bilingual fine-tuning examples, and a 150-item supplied-law control. We evaluate six instruction-tuned models: Llama-3.2-1B, Llama-3.2-3B, Qwen3.5-0.8B, Qwen3.5-2B, Qwen3.5-4B, and Gemma-4-E2B, with three LoRA seeds per model. To separate effects, we combine constrained option-letter scoring, cyclic option rotation, and controlled removal of the governing provision. On 398 Bar Council outputs, an exact-line parser attributes an accuracy gain of 50.0\% to the Qwen3.5-2B seed-42 adapter, whereas option scoring yields only $3.0\%$. For Gemma-4-E2B, the two scoring methods favor different systems. When the governing provision is guaranteed to be present, five of six reference models improve by $14.7\%-19.3\%$ under the four-order criterion. Removing that provision reduces accuracy by $8.0\%-15.3\%$ for models and by $13.8\%-14.9\%$ points for their adapters. However, difference-in differences estimates show no increase in reliance on the governing provision after fine-tuning. Results show that legal adaptation claims require separating scorer, retriever, and model effects. Our Code and data are available at https://anonymous.4open.science/r/bangladesh-legal-qa-11E3
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。