arXiv:2410.09083cs.AIcs.CL2024-10被引 1

检测大模型推理过程中的错误逻辑,即使答案看似正确。

Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment

  • 通过分析输入词元间的交互关系,量化模型推理路径
  • 实验发现多数正确答案背后存在误导性或无关逻辑
  • 适合关注模型可解释性与法律AI可信度的研究者

本文提出一种方法,通过案例研究法律大模型的判断推理模式,依据人类领域知识识别模型潜在的错误表征。不同于传统对生成结果的评估,本研究聚焦于看似正确输出背后的详细推理路径是否正确。我们量化大模型在推理中使用的输入词元之间的交互,因近期理论成果已证明基于交互的解释具有数学上的忠实性保障。为此设计了一套指标来评估大模型的推理模式。实验表明,即使语言生成结果正确,大模型用于法律判断的推理模式中仍有相当比例呈现误导性或无关逻辑。

原文摘要 · Abstract (English)

This paper presents a method to analyze the inference patterns used by Large Language Models (LLMs) for judgment in a case study on legal LLMs, so as to identify potential incorrect representations of the LLM, according to human domain knowledge. Unlike traditional evaluations on language generation results, we propose to evaluate the correctness of the detailed inference patterns of an LLM behind its seemingly correct outputs. To this end, we quantify the interactions between input phrases used by the LLM as primitive inference patterns, because recent theoretical achievements have proven several mathematical guarantees of the faithfulness of the interaction-based explanation. We design a set of metrics to evaluate the detailed inference patterns of LLMs. Experiments show that even when the language generation results appear correct, a significant portion of the inference patterns used by the LLM for the legal judgment may represent misleading or irrelevant logic.

大模型推理可解释性法律AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。