测试大模型在印度法律实务中的表现,发现能写文书却常编造事实。
Evaluating the Role of Large Language Models in Legal Practice in India
- 用真人律师和学生评分,对比GPT、Claude等模型在法律任务上的表现。
- 模型在文书起草和问题识别上接近甚至超过人类,但在专业研究中错误率高。
- 适合法律从业者参考,但复杂案件仍需人类专家判断。
人工智能融入法律行业引发对大语言模型(LLM)执行关键法律任务能力的质疑。本文通过调查实验,实证评估GPT、Claude、Llama等模型在印度法律语境下执行问题识别、法律起草、法律建议、研究与推理等任务的表现。由初级律师和高年级法学生对模型输出进行评分,评价标准包括实用性、准确性和全面性。结果显示,模型在起草和问题识别方面表现优异,常达到或超过人类水平;但在专业法律研究中频繁出现幻觉,生成事实错误或虚构内容。结论是:尽管大模型可辅助部分法律工作,但复杂推理和法律精确适用仍需人类专家参与。
原文摘要 · Abstract (English)
The integration of Artificial Intelligence(AI) into the legal profession raises significant questions about the capacity of Large Language Models(LLM) to perform key legal tasks. In this paper, I empirically evaluate how well LLMs, such as GPT, Claude, and Llama, perform key legal tasks in the Indian context, including issue spotting, legal drafting, advice, research, and reasoning. Through a survey experiment, I compare outputs from LLMs with those of a junior lawyer, with advanced law students rating the work on helpfulness, accuracy, and comprehensiveness. LLMs excel in drafting and issue spotting, often matching or surpassing human work. However, they struggle with specialised legal research, frequently generating hallucinations, factually incorrect or fabricated outputs. I conclude that while LLMs can augment certain legal tasks, human expertise remains essential for nuanced reasoning and the precise application of law.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。