用数学空间理论重新定义AI智能评估,揭示评分体系的内在结构
Psychometric Tests for AI Agents and Their Moduli Space
- 构建心理测评的模空间框架,为智能评分提供数学基础
- 证明已有AAI指数是该框架下的特殊案例,理论统一性更强
- 提出认知核心概念,适合研究智能评估机制的学者参考
我们从模论视角构建AI代理心理测评体系的数学框架,并与先前提出的AAI分数建立明确联系。首先,精确定义了测评电池上的AAI泛函,并提出合理自主性/通用智能评分应满足的公理。其次,证明此前定义的复合指标('AAI-Index')是该AAI泛函的特例。第三,引入相对于测评电池的代理认知核心概念,并将对应的AAI$_{\textrm{core}}$分数定义为AAI泛函在该核心上的限制。最后,利用这些概念描述在保持评估不变的对称性下测评体系的不变量,并概述等价测评体系的模空间组织方式。
原文摘要 · Abstract (English)
We develop a moduli-theoretic view of psychometric test batteries for AI agents and connect it explicitly to the AAI score developed previously. First, we make precise the notion of an AAI functional on a battery and set out axioms that any reasonable autonomy/general intelligence score should satisfy. Second, we show that the composite index ('AAI-Index') defined previously is a special case of our AAI functional. Third, we introduce the notion of a cognitive core of an agent relative to a battery and define the associated AAI$_{\textrm{core}}$ score as the restriction of an AAI functional to that core. Finally, we use these notions to describe invariants of batteries under evaluation-preserving symmetries and outline how moduli of equivalent batteries are organized.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。