揭示掩码预测的识别盲区:模型可能误判数据分布却仍表现良好
An Identifiability Theory of Masked Prediction: Mode Blindness and Mask Schedules
- 通过ε可识别性模量分析掩码预测对联合分布的还原能力
- 在慢混合数据中,模型可对全局模式赋予错误概率而风险仅指数级小
- 低可见性掩码能恢复对模式频率的敏感性,适合研究生成模型泛化
掩码预测通过可见上下文推断缺失变量来学习。这引出一个根本问题:何时近似最优的条件预测能确定联合数据分布?本文通过ε-可识别性模量衡量在掩码预测超出风险不超过ε时,允许的最大联合分布误差。对于具有分离全局模式的慢混合数据,我们发现模型可在不引发显著额外风险的情况下,为整个数据域分配明显错误的概率。精确信息分解表明:固定掩码下,预测损失仅检测可见上下文未解决的部分模式权重偏差;对小的模式权重扰动,该敏感性与剩余模式不确定性成正比。当对掩码平均后,这种剩余不确定性决定了目标函数对全局模式频率的敏感度——低可见性掩码可恢复模式权重敏感性,而全掩码正质量则在联合条件目标下提供对联合分布的普遍控制。我们通过精确计算、受控优化实验及自然文本测量提供了计算与实证支持。更广泛地,本研究表明,预测目标只能识别那些其条件结构未完全解决的全局差异。
原文摘要 · Abstract (English)
Masked prediction learns by inferring missing variables from visible context. This raises a fundamental question: when does near-optimal conditional prediction determine the joint data law? We study this via an $\varepsilon$-identifiability modulus measuring the largest joint-law error compatible with masked-prediction excess risk at most $\varepsilon$. For slow-mixing data laws with separated global modes, we show that a model can assign substantially incorrect probabilities to entire data regimes while incurring exponentially small excess risk. An exact information decomposition reveals why: for a fixed mask, the prediction loss detects only the portion of the mode-weight mismatch that the visible context leaves unresolved. For small mode-weight perturbations, this sensitivity is proportional to residual mode uncertainty. Once averaged over masks, this residual uncertainty governs the objective's sensitivity to global mode frequencies, with low-visibility masks restoring mode-weight sensitivity and positive full-mask mass providing universal joint-law control under the joint conditional objective. We provide computational and empirical evidence for these predictions through exact calculations, controlled optimization experiments, and measurements on natural text. More broadly, our study suggests that a predictive objective can identify global distinctions only insofar as its conditioning structure leaves them unresolved.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。