用上下文一致性检测大模型答案是否可信,发现稳定回答更可能正确。
Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility

- 通过对比原始与扰动提示下的模型输出,衡量答案一致性。
- 一致性高的答案在多个任务中更可能正确或事实无误。
- 适合评估模型可靠性,尤其在评测指标饱和时仍有效。
大型语言模型是强大的黑箱系统,难以判断其回答反映的是稳定的内在信念,还是表面的模式匹配。我们提出将跨上下文一致性(C3)作为未被充分利用的行为特征:可信的回答在主题一致但内容中立的上下文变化下应保持稳定。基于此,我们通过比较模型在原始提示和扰动提示下的生成结果来操作化C3。在26个模型和六个涵盖推理、事实性及代码生成的基准测试中,我们发现跨上下文差异小的答案更可能正确或真实。结果表明,C3提供了互补的评估维度,可作为基准有效性诊断工具,识别出在整体得分趋于饱和时仍具信息量的评测部分。
原文摘要 · Abstract (English)
Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral property of LLMs: a credible answer should remain stable when the same task is placed under topic-aligned, content-neutral contextual variation. Building on this intuition, we operationalize Cross-Contextual Consistency (C3) by comparing model generations under original and perturbed prompts. Across 26 models and six benchmarks spanning reasoning, factuality, and code generation, we find that answers with smaller cross-contextual shifts are more likely to be correct or factual. We demonstrate that C3 provides a complementary axis of evaluation and can serve as a benchmark usefulness diagnostic, identifying which portions of a benchmark remain informative even when aggregated scores are widely considered "saturate".
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。