arXiv:2608.29352cs.AI2026-08

通过建模指令间关系,让大模型更好理解复杂指令的细微差别。

Cross-Relational Preference Learning for Better LLM Instruction Following

论文配图:Cross-Relational Preference Learning for Better LLM Instruction Following
图 1 · 摘自论文原文
  • 用跨指令扰动和区域配对方法生成更丰富的偏好数据
  • 在多个基准上显著提升模型指令遵循能力,效果优于现有方法
  • 适合希望提升模型对复杂指令理解力的研究者与开发者

大型语言模型(LLMs)在遵循复杂指令方面仍存在局限。现有方法虽多采用偏好学习提升该能力,但通常忽略不同指令允许响应空间之间的关系,导致模型难以适应细微且多样的约束变化。为此,我们提出交叉关系偏好学习(CRPL),一种通过两种关键技术构建偏好数据的新框架:交叉关系扰动与跨区域配对采样。该方法能生成更具多样性的偏好数据,充分捕捉各类约束变化。此外,引入基于原子约束的验证机制,严格评估响应满足度,确保偏好对质量。在多种偏好学习方法(如DPO、KTO)、不同模型架构及四个指令遵循基准上的大量实验表明,本方法显著优于先前基线,并展现出强大泛化能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) still exhibit limited capability in following complex instructions. While existing approaches often rely on preference learning to enhance this ability, they typically overlook the relationships between the permissible response spaces of different instructions, which restricts a model to align with subtle and diverse constraint variations. To address this, we propose Cross-Relational Preference Learning (CRPL), a novel framework for constructing preference data that explicitly models inter-instruction relationships through two key techniques: Cross-Relationship Perturbation and Cross-Region Pair Sampling. This enables the generation of more diverse preference data that captures a wide spectrum of constraint variations. Additionally, we introduce an atomic constraint-based verification mechanism to rigorously assess response satisfaction, ensuring high-quality preference pair construction. Extensive experiments across multiple preference learning methods (e.g., DPO, KTO), LLM backbones and four instruction-following benchmarks demonstrate that our approach achieves substantial improvements over prior baselines and exhibits strong generalization.

指令遵循偏好学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。