大模型对复杂法律问题回答不稳定,同一问题多次回答可能矛盾。
LLMs Provide Unstable Answers to Legal Questions
- 用真实案例提炼500个法律问题,测试大模型一致性
- 即使温度设为0,GPT-4o、Claude-3.5等仍对同一问题给出不同判决
- 适用于关注法律AI可靠性的人士,如律师与司法科技开发者
大模型在面对困难法律问题时若多次重复相同提问仍得出不一致结论,则视为不稳定。我们发现,即使将温度设置为0以实现最大程度确定性,GPT-4o、Claude-3.5和Gemini-1.5等主流大模型在回答复杂法律问题时仍表现出显著不稳定性。为此,我们构建并发布了一个包含500个法律问题的新数据集,这些问题源自真实案件,涉及双方当事人、具体事实、相互竞争的法律主张以及‘哪方应胜诉’的核心问题。当输入完全相同的提问时,大模型有时判定一方胜诉,有时又判定另一方胜诉。这一不稳定性对日益增长的法律AI产品、法律流程及依赖这些模型的律师带来了深远影响。
原文摘要 · Abstract (English)
An LLM is stable if it reaches the same conclusion when asked the identical question multiple times. We find leading LLMs like gpt-4o, claude-3.5, and gemini-1.5 are unstable when providing answers to hard legal questions, even when made as deterministic as possible by setting temperature to 0. We curate and release a novel dataset of 500 legal questions distilled from real cases, involving two parties, with facts, competing legal arguments, and the question of which party should prevail. When provided the exact same question, we observe that LLMs sometimes say one party should win, while other times saying the other party should win. This instability has implications for the increasing numbers of legal AI products, legal processes, and lawyers relying on these LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。