实验证明:结构代理无法可靠预测模型在分布外任务上的表现。
A Controlled Counterexample to Strong Proxy-Based Explanations of OOD Performance: in a Fixed Pretraining-and-Probing Setup

- 构建可控反例,分离结构代理与任务相关结构
- 三个种子中两个出现代理排名与真实性能排名相反
- 适用于关注模型解释可信度的研究者
任务无关的结构代理常被用来解释不同预训练语料库的迁移效果差异,但其有效性依赖于代理是否能捕捉下游任务相关的结构。本文在固定预训练与探针设置下检验这一假设,引入计算受限的结构概念(如epiplexity)。核心问题是:结构代理对两个预训练数据集的排序是否应与它们在分布外探针中的准确率排序一致?我们证明并非如此。首先,在形式化构造中,总学习结构量、其操作代理和目标任务相关结构完全分离。随后在合成序列模型实验中,主评估下三个种子中有两个出现分布外准确率排名反转代理排名。辅助诊断与消融分析支持该解释。反例不否定结构解释本身,而是揭示强代理解释的边界:即使在控制环境下,总结构代理也可能无法追踪驱动分布外性能的任务相关结构。
原文摘要 · Abstract (English)
Task-agnostic structure proxies are often used to interpret why one pretraining corpus transfers better than another, but such explanations require the proxy to track the structure that matters for the downstream task. We test this requirement in a fixed pretraining-and-probing setup motivated by computationally bounded notions of learned structure, including epiplexity. The core question is whether a proxy ranking of two pretraining datasets must agree with their ranking by OOD probe accuracy. We show that it need not. First, we give a controlled construction in which a formal structure quantity, its operational proxy, and the task-relevant structure for a target family separate. We then instantiate the same mechanism in a synthetic sequence-model experiment: under the primary all-sample evaluation, the OOD accuracy ranking reverses the proxy ranking in two of three seeds, with auxiliary diagnostics and ablations supporting the same interpretation. The counterexample does not reject structure-based explanations in general; it identifies a boundary on strong proxy-based explanations. A proxy for total learned structure can fail to track the task-relevant structure that drives OOD performance, even in a controlled setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。