arXiv:2609.06190cs.CV2026-09

单次扰动无法验证AI是否使用输入,需至少n次独立扰动才能准确评估。

One Perturbation Is Not Enough: Identifiability and Blind Baselines for Behavioral AI Evaluation

论文配图:One Perturbation Is Not Enough: Identifiability and Blind Baselines for Behavioral AI Evaluation
图 1 · 摘自论文原文
  • 用多组独立扰动的对数来覆盖输入空间,才能精确识别输入依赖关系。
  • 3个视觉-语言模型在测试中得分均低于盲基线,说明其默认使用固定校准值。
  • 适用于评估模型行为可靠性,尤其关注输入使用机制的研究者。

行为评估通过扰动输入并观测输出变化来验证系统是否使用该输入。我们证明,此类认证所需的扰动数量是固定的,仅报告单一扰动无法提供足够证据。当响应比是策略的属性而非测试项的特性时,行为记录是对输入依赖指数向量的线性测量;只有当扰动的对数能张成输入空间时,才能精确识别输入使用情况。对于n个输入,至少需要n个扰动。不完整的扰动设计会混淆在设计矩阵核空间中不同的策略,且强化现有扰动无法替代增加独立扰动。我们还推导出完全不读取输入的策略所获得的评分闭式解,该值远高于零,而我们调查的所有探针均未报告此结果。在正确响应由量纲分析确定的前提下,对三个视觉-语言模型施加完整的三扰动集,要求其从视频中报告物理量。结果显示,所有模型得分均显著低于自身盲基线,因为它们均默认采用少数几个整数校准值,与尺度一致的真值不符;无一模型能通过单一指数级变化改变重标注意图,且均无法与被指令忽略视频的同一模型区分开。

原文摘要 · Abstract (English)

Behavioral evaluations perturb an input and read the induced change in the output in order to certify that a system uses that input. We show that the number of perturbations such a certificate requires is fixed, and that reporting a single perturbation cannot supply it. Where a response ratio is a property of the policy rather than of the test items, the behavioral record is a linear measurement of an exponent vector recording how much the output depends on each input, so perturbations identify input use exactly when their logarithms span the input space. At least $n$ are needed for $n$ inputs, an incomplete design confuses precisely the policies differing along the kernel of its design matrix, and sharpening a perturbation never substitutes for adding an independent one. We also derive in closed form the score such a test awards a policy that reads nothing, which is far from zero and which none of the probes we survey reports. Instantiating this where the correct response is fixed by dimensional analysis, we run a complete identifying set of three perturbations on three vision--language models reporting a physical quantity from video. All three score far below their own blind bound rather than above it, because each defaults to one of a small set of round calibration values that never matches what the scale asserts; none moves its relabeling response by a single exponent, and none is separable from the same model instructed to ignore the video.

AI评估行为分析模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。