评测4大AI模型在孟加拉法律场景下的可靠性,发现其回答质量参差且存严重误导风险。
Assessing the Reliability of Large Language Models in the Bengali Legal Context: A Comparative Evaluation Using LLM-as-Judge and Legal Experts
- 用专家和AI双评估框架,对比4个大模型在真实法律问题上的表现。
- 模型常生成结构良好但含虚假案例引用、错误程序等危险错误信息。
- 适合关注AI司法应用风险的政策制定者与法律科技开发者参考。
在孟加拉,获取法律帮助困难重重:律师费用高、法律语言复杂、律师短缺,且有数百万未结案件。生成式AI模型如OpenAI GPT-4.1 Mini、Gemini 2.0 Flash、Meta Llama 3 70B和DeepSeek R1可能通过提供快速、低成本的法律建议,推动法律服务民主化。本研究从Facebook群组“Know Your Rights”收集了250个真实的法律问题,由四位认证的孟加拉国法律专业人士提供权威解答。所有问题均使用统一提示词提交给四个先进AI模型生成回复。采用双评估框架:一个先进的LLM作为裁判,从事实准确性、法律适当性、完整性与清晰度四方面评估;同时由三位持证孟加拉法律专业人士按相同标准评分。此外,还使用BLEU等自动指标评估回复相似性。结果显示,尽管部分模型能生成高质量、结构良好的回应,但普遍存在严重错误,包括虚构案例引用、错误法律程序及潜在有害建议。研究强调,在孟加拉正式部署AI法律咨询前,必须进行严格专家验证与全面安全防护。
原文摘要 · Abstract (English)
Accessing legal help in Bangladesh is hard. People face high fees, complex legal language, a shortage of lawyers, and millions of unresolved court cases. Generative AI models like OpenAI GPT-4.1 Mini, Gemini 2.0 Flash, Meta Llama 3 70B, and DeepSeek R1 could potentially democratize legal assistance by providing quick and affordable legal advice. In this study, we collected 250 authentic legal questions from the Facebook group "Know Your Rights," where verified legal experts regularly provide authoritative answers. These questions were subsequently submitted to four four advanced AI models and responses were generated using a consistent, standardized prompt. A comprehensive dual evaluation framework was employed, in which a state-of-the-art LLM model served as a judge, assessing each AI-generated response across four critical dimensions: factual accuracy, legal appropriateness, completeness, and clarity. Following this, the same set of questions was evaluated by three licensed Bangladeshi legal professionals according to the same criteria. In addition, automated evaluation metrics, including BLEU scores, were applied to assess response similarity. Our findings reveal a complex landscape where AI models frequently generate high-quality, well-structured legal responses but also produce dangerous misinformation, including fabricated case citations, incorrect legal procedures, and potentially harmful advice. These results underscore the critical need for rigorous expert validation and comprehensive safeguards before AI systems can be safely deployed for legal consultation in Bangladesh.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。