提出双维度难度空间,让模型自适应选择学习路径。
Dual-Difficulty Curriculum Learning for Direct Preference Optimization
- 用提示复杂度和可区分性构建双维难度空间
- 自适应学习使性能超越基线,数据效率提升显著
- 适合需要高效对齐的LLM研究者与开发者
课程学习能提升大语言模型对齐效果,但现有方法仅基于单一难度维度。本文将对齐难度重新定义为由提示复杂度(PC)和成对可区分性(PD)构成的二维空间,提供更坚实的理论基础。首先提出静态课程框架DM-Curri-DPO,已实现显著性能提升。进一步提出核心贡献GSP-Curri-DPO——一种分组自适应学习框架,使模型能根据自身能力动态探索最优学习路径。大量实验表明,该方法不仅在关键基准上达到新最优,还展现出更强的数据效率与抗偏好噪声能力。本工作建立了大模型对齐的新范式,兼具结构化难度空间与智能模型驱动的学习策略。
原文摘要 · Abstract (English)
Curriculum learning enhances Direct Preference Optimization (DPO) for aligning Large Language Models (LLMs), yet existing methods rely on a one-dimensional view of difficulty. In this work, we reframe alignment difficulty as a two-dimensional space spanned by Prompt Complexity (PC) and Pairwise Distinguishability (PD), providing a more principled foundation for alignment. We first demonstrate the efficacy of this space by developing DM-Curri-DPO, a framework of static curricula that already achieves significant gains over baseline methods. Moving beyond these handcrafted paths, we introduce our primary contribution: GSP-Curri-DPO, a novel Group-wise Self-Paced Learning framework. This advanced method empowers the model to navigate the difficulty grid, discovering an optimal learning trajectory based on its own evolving capabilities. Extensive experiments show our self-paced approach not only sets a new state-of-the-art on key benchmarks but, more importantly, demonstrates superior data efficiency and robustness to preference noise. Our work establishes a new paradigm for LLM alignment, offering both a structured difficulty space and an intelligent, model-driven methodology for navigating it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。