arXiv:2508.16357cs.CLcs.AI2025-08被引 4

构建首个摩洛哥法律问答基准,评估大模型在复杂法律语境中的表现

MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering

  • 基于摩洛哥法律体系设计多源融合的问答数据集
  • 超1700道题揭示大模型在法律推理上的显著短板
  • 适合法律AI研究者与跨文化语料开发人员参考

大语言模型(LLM)在自然语言处理领域进展迅速,但在阿拉伯语法律等专业化、低资源领域仍面临挑战。本文提出MizanQA(发音为Mizan,意为阿拉伯语中的“天平”,象征公正),一个针对摩洛哥法律问答任务的基准测试数据集,涵盖现代标准阿拉伯语、伊斯兰马里克法学、摩洛哥习惯法及法国法律影响,具有丰富的语言与法律复杂性。数据集包含超过1700个多项选择题,支持多答案形式,真实反映法律推理过程。对多语言及阿拉伯语专用大模型的基准测试显示显著性能差距,凸显了定制化评估指标和文化语境驱动的领域特定模型发展的必要性。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has significantly propelled progress in natural language processing (NLP). However, their effectiveness in specialized, low-resource domains-such as Arabic legal contexts-remains limited. This paper introduces MizanQA (pronounced Mizan, meaning "scale" in Arabic, a universal symbol of justice), a benchmark designed to evaluate LLMs on Moroccan legal question answering (QA) tasks, characterised by rich linguistic and legal complexity. The dataset draws on Modern Standard Arabic, Islamic Maliki jurisprudence, Moroccan customary law, and French legal influences. Comprising over 1,700 multiple-choice questions, including multi-answer formats, MizanQA captures the nuances of authentic legal reasoning. Benchmarking experiments with multilingual and Arabic-focused LLMs reveal substantial performance gaps, highlighting the need for tailored evaluation metrics and culturally grounded, domain-specific LLM development.

法律AI多语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。