arXiv:2510.17426cs.CLcs.AI2025-10ACL被引 10

通过权重插值融合对齐前后模型,同时提升准确率与校准度。

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging

  • 在对齐前后的模型权重间进行线性插值,实现后处理优化。
  • 插值模型在准确率上超越双亲模型,且显著恢复校准能力。
  • 适合追求高可靠性和高性能的模型部署场景。

后训练中的'对齐代价'通常被理解为任务准确率下降,我们发现它还伴随着严重的校准损失,导致模型过度自信、可靠性降低及输出多样性减少。通过一种简单的后处理干预——对齐前后模型权重进行插值,我们有效缓解了这一权衡。关键在于,这并非严格意义上的取舍:插值过程始终能发现帕累托最优解,即在准确率上超过双亲模型的同时,显著恢复对齐过程中丢失的校准性能。本研究证明,简单的模型融合方法可高效缓解对齐代价的全范围影响,生成更强大且更可靠的模型。

原文摘要 · Abstract (English)

The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, making models overconfident, less reliable, and model outputs less diverse. We show that this trade-off can be navigated effectively via a simple post-hoc intervention: interpolating between a model's weights before and after alignment. Crucially, this is not a strict trade-off. We find that the process consistently reveals Pareto-optimal interpolations - models that improve accuracy beyond both parents while substantially recovering the calibration lost during alignment. Our work demonstrates that simple model merging provides a computationally efficient method for mitigating the full scope of the alignment tax, yielding models that are more capable and more reliable.

模型对齐校准优化权重插值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。