arXiv:2505.17656cs.CL2025-05EMNLP被引 20

发现大模型常重复输出错误答案,现有检测方法失效

Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs

  • 提出自一致错误概念,指模型多次生成相同错误内容
  • 大模型规模越大,自一致错误越难被发现,检测率普遍低于30%
  • 用外部验证模型融合隐藏状态,提升跨模型错误识别能力

由于大语言模型(LLMs)常生成看似合理但错误的内容,错误检测对确保真实性愈发关键。然而,现有检测方法常忽略一种我们称之为自一致错误的问题:即模型在多个随机采样中反复生成相同的错误回答。本文首次形式化定义自一致错误,并评估主流检测方法在该问题上的表现。研究发现:(1) 与随模型规模增大而减少的不一致错误不同,自一致错误频率保持稳定甚至上升;(2) 四类检测方法均显著难以识别自一致错误。这些结果揭示了当前检测方法的关键局限。基于自一致错误在不同模型间存在差异的观察,我们提出一种简单有效的跨模型探针方法,通过融合外部验证模型的隐藏状态证据。该方法在三个不同模型家族上显著提升了对自一致错误的检测性能。

原文摘要 · Abstract (English)

As large language models (LLMs) often generate plausible but incorrect content, error detection has become increasingly critical to ensure truthfulness. However, existing detection methods often overlook a critical problem we term as self-consistent error, where LLMs repeatedly generate the same incorrect response across multiple stochastic samples. This work formally defines self-consistent errors and evaluates mainstream detection methods on them. Our investigation reveals two key findings: (1) Unlike inconsistent errors, whose frequency diminishes significantly as the LLM scale increases, the frequency of self-consistent errors remains stable or even increases. (2) All four types of detection methods significantly struggle to detect self-consistent errors. These findings reveal critical limitations in current detection methods and underscore the need for improvement. Motivated by the observation that self-consistent errors often differ across LLMs, we propose a simple but effective cross-model probe method that fuses hidden state evidence from an external verifier LLM. Our method significantly enhances performance on self-consistent errors across three LLM families.

错误检测大模型自一致性验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。