arXiv:2606.20929cs.CLcs.AI2026-06中稿 · the International …

利用大模型内部特征检测法律分类错误,提升系统可靠性

Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification

论文配图:Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification
图 1 · 摘自论文原文
  • 通过分析大模型内部状态识别输出正确性
  • 在保释与违法判定任务中准确发现错误预测
  • 适合关注AI司法应用可靠性的研究者与从业者

大型语言模型(LLMs)在法律领域应用日益广泛。然而,尽管性能优异,其仍易产生错误或幻觉输出,在高风险领域如法律中引发严重可靠性担忧。检测基于LLM系统的回答正确性成为关键挑战。本文探索利用LLM内部特征来检测法律分类任务中的预测正确性。提出方法基于这些内部特征构建下游分类器,以识别错误输出。在两个典型法律分类任务——保释决策预测和法律条文违反预测上进行评估。实验结果表明,LLM的内部特征是检测法律分类错误的可靠指标,可有效提升基于LLM的分类系统可靠性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly being adopted in the legal domain. However, despite their strong performance, LLMs are prone to generating incorrect or hallucinated outputs, raising serious concerns about their reliability in high-stakes domains such as law. Detecting the correctness of responses of LLM-based systems is therefore a critical challenge. In this work, we explore the potential of leveraging internal artifacts of LLM to detect the correctness of their predictions in legal-domain classification tasks. We develop approaches that utilize features derived from these internal artifacts to build downstream classifiers capable of identifying incorrect LLM outputs. We evaluate our approach on two representative legal classification tasks: bail decision prediction and statute violation prediction. Our experimental results demonstrate that LLMs' internal artifacts are reliable indicators for detecting incorrect predictions in legal classification tasks, and can be applied to enhance the reliability of LLM-based classification systems.

法律AI模型可靠性内部特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。