arXiv:2608.21409cs.CYcs.AI2026-08ACL被引 1

LLMs在法律推理中易被权威信息误导,不如医学领域可靠。

Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?

  • 构建医疗与法律对比框架,评估模型对权威信息的判断能力。
  • 法律LLM在引用冲突时过度自信,错误率比医疗场景高3倍以上。
  • 模型越大越易受权威干扰,适合法律实操者警惕其盲信风险。

在医学中,基于稳定生物事实的证据可维持有效性;而在法律中,真理具有地域性、时效性,并依赖权威来源层级。近年来大语言模型(LLMs)在医学执照考试中的成功,催生了其具备同等法律能力的预期。然而,这一类比忽略了关键差异:法律表现更依赖于判断外部权威是否适用、有效且不矛盾,而非单纯推理。本文提出一个四维对比诊断框架(知识召回、依据性、置信度、鲁棒性),应用于新构建的基准数据集(含时间有效性与规范关系)。结果显示,医疗LLM能有效利用验证来源,而法律LLM难以判断检索引用的有用性或误导性,在扰动情境下表现出过高置信度,且对格式等表面线索敏感。模型规模增大加剧此问题,揭示更强指令遵循能力可能伴随更弱的权威抗干扰能力。研究说明LLM将法律视为无结构文本,而非具有约束力的判例;当外部引用与内部知识冲突时,存在过度信任权威但虚假信息的倾向。

原文摘要 · Abstract (English)

In medicine, claims remain valid when supported by empirical evidence grounded in stable biological reality. In law, by contrast, truth is contingent, defined by jurisdiction, temporal validity, and the hierarchy of authoritative sources. The recent success of large language models (LLMs) on medical licensing examinations has encouraged an expectation of comparable legal competence. This analogy, however, obscures a critical distinction between domains. Unlike in medicine, legal performance often depends less on inference than on determining when external authority is applicable, valid, and non-contradictory. We introduce a comparative diagnostic framework evaluating legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness), uncovering a sharp domain asymmetry when applied to a new benchmark that encodes temporal validity and normative relationships. While medical LLMs reliably benefit from verified sources, legal LLMs struggle to assess when retrieved citations are useful or misleading, exhibiting overconfidence in perturbed contexts and sensitivity to superficial formatting cues. Increased model scale amplifies this tendency, revealing that stronger instruction following can coincide with weaker resistance to authoritative perturbations. These findings show that LLMs treat law as unstructured text rather than binding precedent, while revealing a tendency to over-trust authoritative but false information when external references conflict with a model's internal knowledge.

大模型法律推理权威依赖可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。