提出新拒绝准则,提升代理模型引导大模型生成的效果
On the Rejection Criterion for Proxy-based Test-time Alignment
- 用保守置信度构建新拒绝准则,替代原有信心阈值
- 在多个数据集上优于现有方法,显著提升生成质量
- 适合关注测试时对齐与大模型生成优化的研究者
近期工作提出基于小规模对齐模型作为代理,指导大规模基础模型生成。隐式奖励方法会扭曲大模型分布,而引导方法则在大模型不确定时交由小模型决定下一个词的生成。本文首次表明,两种方法均可归约为采样自相似图模型,仅在拒绝准则(或分布)定义上不同。此外,我们指出原有置信度准则因语言模糊性等问题缺乏合理依据。为此,提出基于保守置信度的新拒绝准则。实验表明,该方法在多个数据集上均优于先前方法。
原文摘要 · Abstract (English)
Recent works proposed test-time alignment methods that rely on a small aligned model as a proxy that guides the generation of a larger base (unaligned) model. The implicit reward approach skews the large model distribution, whereas the nudging approach defers the generation of the next token to the small aligned model when the large base one is unconfident about its outcome. In this work, we first show that both approaches can be reduced to sampling from similar graphical models, where they differ only in the definition of a rejection criterion (or distribution). Moreover, we argue that the confidence criterion is ill-motivated due to linguistic phenomena like ambiguous phrasing. We propose a novel rejection criterion based on a conservative confidence bet. Experimentally, our novel approach outperforms previous work on several datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。