提出两阶段框架缓解强模型在弱监督下的过拟合问题
How to Mitigate Overfitting in Weak-to-strong Generalization?
- 分两阶段优化监督信号与输入问题质量
- 在数学基准上实现最高100%的PGR提升
- 适合关注大模型对齐与泛化能力的研究者
在超级对齐(superalignment)任务中,弱到强泛化旨在通过弱监督者激发强模型的能力,确保其行为符合弱监督者的意图且无欺骗等不安全行为。然而,强模型在该过程中易出现过拟合:由于强模型具备强大拟合能力,弱监督者提供的错误标签会导致强模型过拟合。单纯过滤错误标签又会降低问题质量,削弱模型在难题上的泛化能力。为此,本文提出一个两阶段框架,同时提升监督信号质量和输入问题质量。在三组大语言模型和两个数学基准上的实验表明,相比朴素弱到强泛化,本框架显著提升了PGR,部分模型甚至达到100%的PGR。
原文摘要 · Abstract (English)
Aligning powerful AI models on tasks that surpass human evaluation capabilities is the central problem of \textbf{superalignment}. To address this problem, weak-to-strong generalization aims to elicit the capabilities of strong models through weak supervisors and ensure that the behavior of strong models aligns with the intentions of weak supervisors without unsafe behaviors such as deception. Although weak-to-strong generalization exhibiting certain generalization capabilities, strong models exhibit significant overfitting in weak-to-strong generalization: Due to the strong fit ability of strong models, erroneous labels from weak supervisors may lead to overfitting in strong models. In addition, simply filtering out incorrect labels may lead to a degeneration in question quality, resulting in a weak generalization ability of strong models on hard questions. To mitigate overfitting in weak-to-strong generalization, we propose a two-stage framework that simultaneously improves the quality of supervision signals and the quality of input questions. Experimental results in three series of large language models and two mathematical benchmarks demonstrate that our framework significantly improves PGR compared to naive weak-to-strong generalization, even achieving up to 100\% PGR on some models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。