arXiv:2609.04127cs.AI2026-09

为大模型推荐提供可信依据,判断何时该信任具体建议。

Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable

论文配图:Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable
图 1 · 摘自论文原文
  • 提出'认识论依据'概念,量化模型推荐的稳定性与适用范围。
  • 设计四级依赖证书,区分不稳定、依赖上下文、局部支持到广泛支持的推荐。
  • 实验证明其优于口头自信度,适合决策者评估模型建议可靠性。

大语言模型日益用于支持组织决策,但用户常缺乏判断是否应信赖特定建议的可靠依据。现有方法多关注模型整体属性(如可靠性、不确定性)或用户信任,而非具体建议背后的依赖基础。本文借鉴认识论理论,提出‘认识论依据’这一决策层级概念,用于表征模型偏好在多大范围内稳定有效。通过四层依赖证书对成对推荐进行操作化:不稳定、上下文依赖、局部支持、广泛支持。使用已知组测试成功恢复专家预设的依据排序,且更强依据与众包工作者独立共识显著一致。此外,该依据信息独立于口头信心表达,也不可由决策难度解释。该框架为无客观真值时,评估个体模型推荐的可信程度提供了理论扎实且可落地的方法。

原文摘要 · Abstract (English)

Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncertainty, or robustness, or focus on user trust, rather than the underlying basis for relying on an individual recommendation. Adapting theoretical foundations from epistemology, we introduce epistemic warrant, a decision-level construct that characterizes the stability of a model's preference and the scope over which that preference holds. We operationalize this construct through a four-tier reliance certificate for pairwise recommendations, distinguishing among unstable, context-dependent, locally supported, and broadly supported recommendations. We validate the construct using contemporary methodologies: known-groups tests successfully recover expert-prespecified warrant orderings, and stronger warrants systematically align with independent consensus from crowd workers. Furthermore, we demonstrate that epistemic warrant provides information distinct from verbalized confidence and is not readily explained by decision difficulty. Ultimately, this framework offers a theoretically grounded, implementable approach for characterizing the warrant of individual LLM recommendations when objective ground truth is unavailable.

大模型可信推荐认知依据决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。