用混合方法精准定位安大略省租房法条,效果远超单纯微调或检索。
Train, Retrieve, or Both? A Four-Arm Head-to-Head for Correct Statutory Citation on the Ontario Residential Tenancies Act

- 结合微调与检索的混合模型提升法律条文引用准确率。
- 混合模型在真实数据集上达到0.481精确匹配,零幻觉引用。
- 小模型加简单检索即可超越大模型复杂流程,适合资源有限场景。
自代表租户、房东及客服人员需要准确指向解答问题的法律条文及其正确引用。本文针对2006年安大略省住宅租赁法案(RTA)及其核心条例,实证研究:仅微调是否足够,还是必须结合检索?在Qwen2.5-7B-Instruct上进行四组对比实验(零样本基线、仅LoRA微调、仅RAG检索、SFT+RAG混合),在小型且待人工验证的真实评估集上以条款+子节精确匹配率评分。基线模型无法引用RTA,仅微调模型会误引条款;检索可从根本上消除幻觉;混合模型表现最佳,精确匹配率达0.481,无幻觉引用。其优势源于微调使条款选择更稳健,能应对高召回候选集带来的干扰。值得注意的是,此低成本bge-small混合方案性能媲美甚至超越基于更大嵌入模型与交叉编码重排序器的复杂流水线,且更大/改进训练集亦无提升。强法定引用能力无需专用检索模型或更多数据。结果虽成功压倒基线并消除幻觉,但尚未达到0.70的理想目标。所有结果基于小规模真实评估集,为初步结论。
原文摘要 · Abstract (English)
Self-represented tenants, landlords, and help-desk staff need to be pointed at the provision of law that actually governs a question, with a correct statutory citation. We study this task on the Ontario Residential Tenancies Act, 2006 (RTA) and its core regulation, asking the operator's question empirically: is fine-tuning enough, or is hybrid retrieval needed? We run a four-arm head-to-head on Qwen2.5-7B-Instruct (base zero-shot, LoRA SFT-only, RAG-only, and an SFT+RAG hybrid), scored on citation exact-match (section+subsection) over a small, human-verification-pending real eval set. The base model cannot cite the RTA and SFT-only mis-recalls sections; retrieval is essential and drives hallucination to zero by construction; and the SFT+RAG hybrid scores highest at 0.481 exact-match with zero hallucinated citations. Its edge comes from SFT making provision selection more robust to the higher-recall candidate sets that hurt zero-shot RAG. Notably, this cheap bge-small hybrid matches or beats a pipeline built on bigger, specialized retrieval models (a larger embedder and a cross-encoder reranker), and a larger/improved training set does not help either: strong statutory-citation performance here does not require specialized retrieval models or more data. The artifact zeroes hallucination and clears the lift-over-base bar but does not reach the aspirational 0.70 exact-match target. All results are on a small, human-verification-pending real eval set and are reported as preliminary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。