arXiv:2508.08466cs.CL2025-08被引 1

轻量级优化让小模型更懂人类偏好,资源有限时效果更优。

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints

  • 基于DPO改进,引入动态阈值和重点样本优化机制。
  • 在AlpacaEval上胜率提升2.0点,长文本任务提升1.4点。
  • 适合算力受限场景下的小模型对齐优化,部署友好。

小型大语言模型在严重性能差距下难以对齐人类偏好。本文提出两种轻量级DPO变体——自适应边际-逻辑斯蒂损失与APO-hinge-zero,通过引入边际目标和选择性更新机制,更好应对表现不足场景。APO-hinge-zero结合铰链损失的困难样本挖掘与APO-zero的选择性优化,在AlpacaEval上相比APO-zero基准提升2.0点胜率、1.4点长度控制胜率;在MT-Bench上各类任务表现优异,尤其在STEM与人文学科任务中突出。结果表明,仅通过偏好目标的简单修改,即可显著提升资源受限下小模型的对齐能力,为高效部署提供可行路径。

原文摘要 · Abstract (English)

Small large language models (LLMs) often face difficulties in aligning output to human preferences, particularly when operating under severe performance gaps. In this work, we propose two lightweight DPO-based variants -- Adaptive Margin-Sigmoid Loss and APO-hinge-zero -- to better address underperformance scenarios by introducing margin-based objectives and selective update mechanisms. Our APO-hinge-zero method, which combines hinge-induced hard-example mining with the chosen-focused optimization of APO-zero, achieves strong results. In AlpacaEval, APO-hinge-zero improves the win rate by +2.0 points and the length-controlled win rate by +1.4 points compared to the APO-zero baseline. In MT-Bench, our methods maintain competitive performance in diverse categories, particularly excelling in STEM and Humanities tasks. These results demonstrate that simple modifications to preference-based objectives can significantly enhance small LLM alignment under resource constraints, offering a practical path toward more efficient deployment.

小模型对齐轻量优化资源约束奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。