arXiv:2602.20400cs.LGcs.AI2026-02

提出三大真实数据挑战,警示当前无监督对齐方法的可靠性危机

Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation

  • 构建三类缺失真实数据特性的测试集,模拟现实场景复杂性
  • 现有无监督对齐方法在新数据上表现显著下降,无一能稳定应对挑战
  • 适合关注大模型安全对齐、实证评估可信度的研究者阅读

为引导语言模型在人类无法判断的任务中输出真实结果,已有研究提出通过简单任务训练来提升复杂任务表现(易到难泛化),或完全无需外部标注的无监督对齐方法。尽管这些技术在广泛任务上提升了模型准确率,我们指出其评估所用数据集存在偏差:通常不具备比真实性更显著的特征、训练集平衡、且仅包含模型可给出明确答案的数据点。为此,我们构造了缺失上述任一特性的数据集,用于压力测试主流无监督对齐与易到难泛化方法。结果表明,没有任何一种技术能在所有挑战下保持可靠表现。同时,集成与混合策略仅部分缓解性能退化。我们认为,克服这些挑战应成为未来无监督对齐研究的优先方向。

原文摘要 · Abstract (English)

To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones (easy-to-hard generalization), or using unsupervised training algorithms to steer models with no external labels at all (unsupervised elicitation). Although techniques from both paradigms have been shown to improve model accuracy on a wide variety of tasks, we argue that the datasets used for these evaluations could cause overoptimistic evaluation results. Unlike many real-world datasets, they often (1) have no features with more salience than truthfulness, (2) have balanced training sets, and (3) contain only data points to which the model can give a well-defined answer. We construct datasets that lack each of these properties to stress-test a range of standard unsupervised elicitation and easy-to-hard generalization techniques. We find that no technique reliably performs well on any of these challenges. We also study ensembling and combining easy-to-hard and unsupervised techniques, and find they only partially mitigate performance degradation due to these challenges. We believe that overcoming these challenges should be a priority for future work on unsupervised elicitation.

模型安全无监督对齐评估漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。