构建巴西法律考试专用评估基准,测试AI模型表现并强调本地化数据重要性
LegalScore: Development of a Benchmark for Evaluating AI Models in Legal Career Exams in Brazil
- 建立LegalScore基准,评估14类AI模型在巴西法律考试中的答题表现
- 开源与本地模型因训练数据更贴合巴西法律语境,表现优于通用模型
- 适用于关注AI在法律教育、考试研发及本地化模型开发的研究者
本研究提出LegalScore,一个专门用于评估生成式人工智能模型在巴西特定职业法律考试中表现的基准。该基准对14种不同类型的AI模型(包括专有和开源模型)在回答客观题方面的表现进行评估。研究发现,使用英语训练的大语言模型在应对巴西法律情境时表现受限,凸显了开发巴西本地化训练数据的重要性。尽管专有和主流模型整体表现更优,但部分本地化和小型模型因训练数据更契合巴西法律环境,展现出令人期待的潜力。通过准确率、置信区间和标准化评分等指标,LegalScore建立了系统化评估框架,支持对AI在巴西法律考试中应用的持续研究。研究认为,尽管AI在备考和题目设计上具有潜力,但在高级法律评估中仍远未达到人类水平。该基准为未来研究奠定了基础,强调了人工智能本地化适配的关键意义。
原文摘要 · Abstract (English)
This research introduces LegalScore, a specialized index for assessing how generative artificial intelligence models perform in a selected range of career exams that require a legal background in Brazil. The index evaluates fourteen different types of artificial intelligence models' performance, from proprietary to open-source models, in answering objective questions applied to these exams. The research uncovers the response of the models when applying English-trained large language models to Brazilian legal contexts, leading us to reflect on the importance and the need for Brazil-specific training data in generative artificial intelligence models. Performance analysis shows that while proprietary and most known models achieved better results overall, local and smaller models indicated promising performances due to their Brazilian context alignment in training. By establishing an evaluation framework with metrics including accuracy, confidence intervals, and normalized scoring, LegalScore enables systematic assessment of artificial intelligence performance in legal examinations in Brazil. While the study demonstrates artificial intelligence's potential value for exam preparation and question development, it concludes that significant improvements are needed before AI can match human performance in advanced legal assessments. The benchmark creates a foundation for continued research, highlighting the importance of local adaptation in artificial intelligence development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。