arXiv:2509.07955cs.LGcs.AI2025-09

通过选择性分歧提升模型在完全伪相关下的泛化能力。

ACE and Diverse Generalization via Selective Disagreement

  • 基于自训练框架,鼓励模型对新样本产生自信且有选择性的分歧预测。
  • 在多个完全伪相关基准上表现优于或匹配现有方法,且对不完全伪相关也鲁棒。
  • 支持先验知识注入和无监督模型选择,适合需可解释决策的场景。

深度神经网络极易受伪相关影响——即模型学习到仅在训练数据中成立的捷径,导致分布外性能下降。现有研究多针对不完整伪相关,依赖可打破相关性的标注样本;但当伪相关为完全时,正确泛化本质上是未定义的。为此,我们提出一种方法,学习一组与训练数据一致、但在部分新未标注输入上做出不同预测的概念。采用自训练策略,鼓励模型在预测中表现出自信且选择性的分歧。该方法名为ACE,在一系列完全伪相关基准上表现匹配或超越现有方法,同时对不完全伪相关仍保持鲁棒性。此外,ACE更易配置,支持先验知识编码和合理的无监督模型选择。在语言模型对齐的早期应用中,即使无未受信任测量数据,也能在测量篡改检测任务上取得竞争力表现。尽管仍有局限,但显著推进了对未定义泛化的应对。

原文摘要 · Abstract (English)

Deep neural networks are notoriously sensitive to spurious correlations - where a model learns a shortcut that fails out-of-distribution. Existing work on spurious correlations has often focused on incomplete correlations,leveraging access to labeled instances that break the correlation. But in cases where the spurious correlations are complete, the correct generalization is fundamentally \textit{underspecified}. To resolve this underspecification, we propose learning a set of concepts that are consistent with training data but make distinct predictions on a subset of novel unlabeled inputs. Using a self-training approach that encourages \textit{confident} and \textit{selective} disagreement, our method ACE matches or outperforms existing methods on a suite of complete-spurious correlation benchmarks, while remaining robust to incomplete spurious correlations. ACE is also more configurable than prior approaches, allowing for straight-forward encoding of prior knowledge and principled unsupervised model selection. In an early application to language-model alignment, we find that ACE achieves competitive performance on the measurement tampering detection benchmark \textit{without} access to untrusted measurements. While still subject to important limitations, ACE represents significant progress towards overcoming underspecification.

模型鲁棒性伪相关自训练泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。