验证器能提升法律问答模型表现,尤其在计算资源有限时
Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks
- 用7个评分模型测试验证器在法律多选题中的效果
- 低预算下过程验证比结果验证更有效,提升约12%准确率
- 小模型配专业验证器效果最佳,适合法律AI研发者
测试时缩放(TTS)技术可在增加计算开销的前提下提升大语言模型(LLMs)性能。尽管该技术在数学和编程等正式领域已证明有效,但在法律等论证性领域仍缺乏充分研究。本文针对五个基准数据集上的法律多选题问答任务,实证评估了基于验证器的TTS方法。采用7个奖励模型,在真实低$N$预算条件下比较结果级(Best-of-$N$)与过程级(树搜索)验证的效果。系统分析表明,验证器效用受领域专属性、模型规模及监督类型(过程监督的PRM vs. 仅结果的ORM)显著影响,且跨角色应用时仍保持有效性。
原文摘要 · Abstract (English)
Test-time scaling (TTS) techniques can improve the performance of large language models (LLMs) at the expense of additional computation and latency. While TTS has proven effective in formal domains such as mathematics and programming, its value in argumentative domains such as law remains underexplored. We present an empirical study of verifier-based TTS methods for legal multiple-choice QA (MCQA) across five benchmarks. Using a family of 7 reward models, we evaluate both outcome-level (Best-of-$N$) and process-level (tree search) verification under realistic low-$N$ budgets. Our analysis systematically investigates how verifier utility is affected by key properties such as domain specialization, model size, and supervision type (process-supervised PRMs vs. outcome-only ORMs), even when applied across different roles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。