arXiv:2608.05015econ.THcs.AI2026-08被引 1

用数学定理无标签评估大模型决策合理性,不依赖人工标注。

Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems

  • 基于决策理论定理,从模型自答中检验行为是否符合理性准则。
  • 三种理性标准下惩罚值连续且为零时行为可被完全合理化。
  • 适合用于无监督评估与训练,补充现有评价信号。

决策理论中的表示定理表明,行为满足特定公理当且仅当其可由明确目标解释。本文提出利用这一“当且仅当”结构,为大模型等AI系统提供无标签评估与正则化的新范式。通过设计合成选择问题,仅需模型自身响应即可检验公理符合性,无需外部标签或人类反馈,且惩罚项可直接计算。由于公理具有必要性和充分性,所构建的检验能穷尽相关理性标准在所采集数据下的所有推论:通过检验的模型无法再被同一组数据以理性理由否定。本文讨论三种实现:de Finetti 定理下的概率一致性、Afriat 定理下的偏好理性,以及 Echenique 与 Saito(2015)定理下的主观期望效用,每种均生成连续惩罚项,当行为可被合理化时惩罚为零。由于一致性不约束具体理性目标,此类惩罚与现有评估和训练信号互补而非替代。

原文摘要 · Abstract (English)

Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well-defined objective. I argue that this ``if and only if'' structure provides a potentially useful foundation for label-free evaluation and regularization of LLMs and other AI systems. Axiom compliance can be checked from the model's own responses to synthetic choice problems, with no external labels or human feedback, and the penalties are readily computable. Because the axioms are necessary and sufficient, the resulting checks exhaust the implications of the relevant rationality standard for the elicited data: a model that passes cannot be rejected on rationality grounds by any further test of the same data. I discuss three instantiations: probabilistic coherence via a theorem of de Finetti, preference rationality via Afriat's theorem, and subjective expected utility via a theorem of Echenique and Saito (2015), each yielding a continuous penalty that is zero whenever behavior can be rationalized. Since coherence does not restrict which objective rationalizes behavior, these penalties complement rather than replace other evaluation and training signals.

无标签评估理性检验大模型决策理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。