用信息论设计新评估法,让AI自评更抗干扰。
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
- 通过提示测试系统间信息关系,而非直接评分
- 在对抗攻击下,总变差距离仍保持0.70-0.77的评估效能
- 无需真实标签,可实现可靠个体项检测
我们提出一种无需真实标签的AI系统评估方法,基于信息论分析对抗性操纵下的信息损失。该方法将评估者视为策略性玩家,通过提示机制估计互信息,使诚实报告成为最优策略。研究表明,特定f散度(如总变差距离,TVD)在攻击下仍具有多项式保障,突破了最坏情况认证中互信息估计的指数壁垒。在对抗攻击下,TVD-MI评估效果稳定(曲线下面积0.70–0.77),而其他方法性能退化至随机水平。该机制将成对评估分解为无真实标签的可靠个体级检测分数,解决了标准同伴预测的关键局限。预注册地址:https://osf.io/c7pum。
原文摘要 · Abstract (English)
We evaluate artificial intelligence (AI) systems without ground truth by exploiting a link between strategic gaming and information loss. Building on established information theory, we analyze which mechanisms resist adversarial manipulation. This motivates mutual evaluation, where the overseer is treated as a strategic player estimating mutual information by prompting, making truthful agent reporting an optimal strategy. We show that certain f-divergences, such as total variation distance (TVD), maintain polynomial guarantees under attack, building on an established exponential barrier for estimating mutual information (MI) in worst-case certification settings. Under adversarial attacks, TVD-MI maintains effectiveness (area under the curve 0.70--0.77) while other approaches can decay toward chance, demonstrating that prompting the same system for information relationships rather than quality judgments can improve robustness. The mechanisms decompose pairwise evaluations into reliable item-level detection scores without ground truth, addressing a key limitation of standard peer prediction. Pre-registration: https://osf.io/c7pum .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。