arXiv:2507.09638cs.CL2025-07被引 3

用新方法提升泰语法律问答的引用准确率和回答质量

Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?

  • 用分组相对策略优化提升大模型法律推理能力
  • 引用准确率最高提升90%,综合评分提高31%
  • 计算成本降低2.5倍,适合资源有限的法律AI开发

检索增强生成(RAG)系统在泰语法律问答任务中的表现仍受限,尤其在需要复杂法律推理的问题上。为解决这一问题,本文提出基于分组相对策略优化(GRPO)的方法,利用BGE-M3嵌入作为低成本语义相似性奖励信号,相比大型语言模型裁判者可将计算开销降低最多2.5倍。在NitiBench基准上的实验表明,该方法使引用F1得分相比基础模型最高提升90%,联合质量指标较指令微调提升31%。关键的是,该方法在复杂法律推理任务上表现出更强鲁棒性,为提升泰语法律大模型提供了一种高效且资源节约的解决方案。

原文摘要 · Abstract (English)

The Retrieval-Augmented Generation (RAG) systems' performance on Thai legal question answering is still limited, especially for questions requiring extensive, complex legal reasoning. To address these limitations, we introduce an approach aligning LLMs toward improved law citation accuracy and better response quality using Group-Relative Policy Optimization (GRPO). Our approach leverages BGE-M3 embeddings as a cost-efficient semantic-similarity reward, significantly reducing computational expenses up to 2.5x compared to large language model judges. Experiments on the NitiBench benchmark demonstrate substantial improvements: GRPO achieves up to 90% citation-F1 gains from the base model and a 31% increase in joint quality metrics over instruction tuning. Crucially, our method shows enhanced robustness on complex legal reasoning tasks compared to instruction tuning, providing an effective and resource-efficient solution for enhancing Thai legal LLMs.

法律AIRAGGRPO泰语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。