arXiv:2603.04968cs.CLcs.AI2026-03

用弱模型高置信度样本替代人工标注,能更低成本地提升模型对齐效果。

When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger

  • 用弱模型的高置信度输出作为训练样本,替代全量人工标注。
  • 仅用20%人工标注+弱模型置信度加权,性能超过100%人工标注的标准DPO。
  • 适用于多种偏好优化任务,显著降低对人工标注的依赖。

偏好对齐是使大语言模型符合人类价值观的关键步骤,但现有方法通常依赖昂贵的人工标注或大规模API模型。我们探索了弱语言模型是否可作为有效标注者。令人意外的是,仅选取弱模型中高置信度的样本,其性能远超完整的人工标注。基于此,我们提出置信度加权偏好优化(CW-PO)框架,通过弱模型的置信度重加权训练样本,适用于多种偏好优化目标。值得注意的是,使用仅20%人工标注的CW-PO模型,在标准DPO设置下表现优于使用100%人工标注的模型。结果表明,结合置信度加权,弱模型可大幅降低偏好对齐成本,甚至超越完全人工标注的方法。

原文摘要 · Abstract (English)

Preference alignment is an essential step in adapting large language models (LLMs) to human values, but existing approaches typically depend on costly human annotations or large-scale API-based models. We explore whether a weak LLM can instead act as an effective annotator. We surprisingly find that selecting only a subset of a weak LLM's highly confident samples leads to substantially better performance than using full human annotations. Building on this insight, we propose Confidence-Weighted Preference Optimization (CW-PO), a general framework that re-weights training samples by a weak LLM's confidence and can be applied across different preference optimization objectives. Notably, the model aligned by CW-PO with just 20% of human annotations outperforms the model trained with 100% of annotations under standard DPO. These results suggest that weak LLMs, when paired with confidence weighting, can dramatically reduce the cost of preference alignment while even outperforming methods trained on fully human-labeled data.

偏好对齐弱模型置信度加权低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。