arXiv:2508.15854cs.CL2025-08被引 9

用微调+检索增强,让中文模型在伊斯兰继承推理上超越顶级大模型。

QU-NLP at QIAS 2025 Shared Task: A Two-Phase LLM Fine-Tuning and Retrieval-Augmented Generation Approach for Islamic Inheritance Reasoning

  • 用LoRA微调阿拉伯语大模型,再结合检索增强生成。
  • 测试准确率达85.8%,高级推理准确率高达97.6%。
  • 适合研究宗教法律推理或阿拉伯语大模型的开发者参考。

本文介绍了我们在QIAS 2025共享任务中子任务1:伊斯兰继承推理的解决方案。我们基于因果语言模型Fanar-1-9B,采用低秩适配(LoRA)进行微调,并集成到检索增强生成(RAG)流水线中。系统需处理伊斯兰继承法中的复杂问题,包括理解继承场景、识别合格继承人、应用固定份额规则及精确计算。在最终测试中,系统准确率达到0.858,优于GPT 4.5、LLaMA、Fanar、Mistral和ALLaM等模型在零样本提示下的表现。结果表明,QU-NLP达到接近最先进水平(85.8%),尤其在高级推理任务上表现突出(97.6%),超越Gemini 2.5和OpenAI o3。这说明领域微调结合检索增强,可使中等规模阿拉伯语大模型在伊斯兰继承推理上超越前沿模型。

原文摘要 · Abstract (English)

This paper presents our approach and results for SubTask 1: Islamic Inheritance Reasoning at QIAS 2025, a shared task focused on evaluating Large Language Models (LLMs) in understanding and reasoning within Islamic inheritance knowledge. We fine-tuned the Fanar-1-9B causal language model using Low-Rank Adaptation (LoRA) and integrated it into a Retrieval-Augmented Generation (RAG) pipeline. Our system addresses the complexities of Islamic inheritance law, including comprehending inheritance scenarios, identifying eligible heirs, applying fixed-share rules, and performing precise calculations. Our system achieved an accuracy of 0.858 in the final test, outperforming other competitive models such as, GPT 4.5, LLaMA, Fanar, Mistral and ALLaM evaluated with zero-shot prompting. Our results demonstrate that QU-NLP achieves near state-of-the-art accuracy (85.8%), excelling especially on advanced reasoning (97.6%) where it outperforms Gemini 2.5 and OpenAI's o3. This highlights that domain-specific fine-tuning combined with retrieval grounding enables mid-scale Arabic LLMs to surpass frontier models in Islamic inheritance reasoning.

伊斯兰法律大模型微调检索增强阿拉伯语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。