用合成数据微调小模型,让AI读懂巴勒斯坦法律
ALKAFI-LLAMA3: Fine-Tuning LLMs for Precise Legal Understanding in Palestine
- 用巴勒斯坦法律文本生成问答对,微调量化版Llama-3模型
- 在零样本和少样本场景下均表现良好,支持多类型法律问答
- 适合资源有限地区部署,为本地化法律AI提供可行路径
大型语言模型(LLMs)在多个领域展现出巨大潜力,但在法律领域,尤其是低资源背景下应用仍受限。本文针对巴勒斯坦法律领域面临的政局不稳、法律体系碎片化及人工智能资源匮乏等问题,提出基于量化版Llama-3.2-1B-Instruct的微调模型,训练数据源自巴勒斯坦法律文本生成的合成数据集。通过使用小型模型与策略性生成的问答对,实现低成本、本地可持续的法律辅助解决方案,可提供准确且上下文相关的法律建议。实验表明,该模型在各类查询任务中表现良好,涵盖是非问答、叙述解释及复杂法律区分,但对计算类问题和结构化列表格式处理仍有不足。本工作为资源受限环境中的AI法律工具部署提供了可行路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable potential in diverse domains, yet their application in the legal sector, particularly in low-resource contexts, remains limited. This study addresses the challenges of adapting LLMs to the Palestinian legal domain, where political instability, fragmented legal frameworks, and limited AI resources hinder effective machine-learning applications. We present a fine-tuned model based on a quantized version of Llama-3.2-1B-Instruct, trained on a synthetic data set derived from Palestinian legal texts. Using smaller-scale models and strategically generated question-answer pairs, we achieve a cost-effective, locally sustainable solution that provides accurate and contextually relevant legal guidance. Our experiments demonstrate promising performance on various query types, ranging from yes/no questions and narrative explanations to complex legal differentiations, while highlighting areas for improvement, such as handling calculation-based inquiries and structured list formatting. This work provides a pathway for the deployment of AI-driven legal assistance tools tailored to the needs of resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。