arXiv:2601.01532cs.AIcs.CL2026-01

提出新方法量化大模型推理时的信念坚定程度。

Aletheia: Quantifying Cognitive Conviction in Reasoning Models via Regularized Inverse Confusion Matrix

  • 用正则化逆混淆矩阵分析模型信念强度。
  • 发现模型在对抗压力下会过度思考。
  • 引入对齐信念分,确保安全与信念一致。

在迈向通用人工智能(AGI)的进程中,现有评估体系面临认识论危机。静态基准仅衡量知识广度,无法量化信念深度。尽管Simhi等人(2025)定义了标准问答中的CHOKE现象,本文将其扩展至系统2推理模型,提出项目Aletheia——一种基于Tikhonov正则化的认知物理学框架,通过逆推评判者的混淆矩阵来量化“认知坚定性”。为避免依赖不可见私有数据,我们设计合成代理协议进行验证。对2025年基线模型(如DeepSeek-R1、OpenAI o1)的初步研究显示,推理模型虽充当“认知缓冲”,但在对抗压力下可能表现出“防御性过度思考”。此外,我们提出对齐信念分数(S_aligned),证明坚定信念不会损害安全性。本工作为衡量AI科学完整性提供了蓝图。

原文摘要 · Abstract (English)

In the progressive journey toward Artificial General Intelligence (AGI), current evaluation paradigms face an epistemological crisis. Static benchmarks measure knowledge breadth but fail to quantify the depth of belief. While Simhi et al. (2025) defined the CHOKE phenomenon in standard QA, we extend this framework to quantify "Cognitive Conviction" in System 2 reasoning models. We propose Project Aletheia, a cognitive physics framework that employs Tikhonov Regularization to invert the judge's confusion matrix. To validate this methodology without relying on opaque private data, we implement a Synthetic Proxy Protocol. Our preliminary pilot study on 2025 baselines (e.g., DeepSeek-R1, OpenAI o1) suggests that while reasoning models act as a "cognitive buffer," they may exhibit "Defensive OverThinking" under adversarial pressure. Furthermore, we introduce the Aligned Conviction Score (S_aligned) to verify that conviction does not compromise safety. This work serves as a blueprint for measuring AI scientific integrity.

认知建模信念量化推理评估AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。