arXiv:2603.08575cs.AIcs.LG2026-03

用可信度定义信任,让AI可信赖。

Trust via Reputation of Conviction

  • 以独立共识验证的可信度作为信任基础
  • 可信度高者声誉持续积累,错误者自然淘汰
  • 适合评估AI代理的可靠性与透明性

本文通过数学建模探讨知识、真理与信任的关系。将真理定义为可重复感知的知识子集,将信息源视为兼具生成与判别功能的实体,提出基于‘可信度’(conviction)的声誉框架——即某观点被独立共识证实的可能性。认为可信度是比正确性或忠诚度更根本的信任依据:其不依赖特定语境,激励真实贡献,并要求透明自洽的感知以支持外部验证。声誉被形式化为在一系列主张上预期加权有符号可信度之和,分析其在不同源-主张情境下的行为特征,指出持续验证既是理论必需,也是声誉累积的实际机制。该框架应用于人工智能代理,识别出它们虽具能力但易出错,唯有可验证的可信度与持续累积的声誉,才能构成对它们信任的可靠基础。

原文摘要 · Abstract (English)

The question of \emph{knowledge}, \emph{truth} and \emph{trust} is explored via a mathematical formulation of claims and sources. We define truth as the reproducibly perceived subset of knowledge, formalize sources as having both generative and discriminative roles, and develop a framework for reputation grounded in the \emph{conviction} -- the likelihood that a source's stance is vindicated by independent consensus. We argue that conviction, rather than correctness or faithfulness, is the principled basis for trust: it is regime-independent, rewards genuine contribution, and demands the transparent and self-sufficient perceptions that make external verification possible. We formalize reputation as the expected weighted signed conviction over a realm of claims, characterize its behavior across source-claim regimes, and identify continuous verification as both a theoretical necessity and a practical mechanism through which reputation accrues. The framework is applied to AI agents, which are identified as capable but error-prone sources for whom verifiable conviction and continuously accrued reputation constitute the only robust foundation for trust.

信任建模可信度AI代理声誉系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。