arXiv:2505.12845cs.AI2025-05

通过挖掘提示中的深层信号,提升大模型在复杂多指令任务中的对齐能力。

Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks

  • 构建跨样本与内样本双层次偏好数据,捕捉更丰富的语义差异。
  • 在多个基准上显著提升模型对复杂指令的遵循能力,效果优于基线。
  • 适用于需要高精度指令理解的场景,如智能助手、内容创作。

强化学习与人类反馈(RLHF)已成为对齐人工智能系统与人类偏好的主流方法,在指令遵循任务中表现优异;然而在复杂多指令任务中仍存在合规性不足的问题。传统方法依赖人工标注或大型语言模型,带来资源消耗或潜在偏差;而现有的合成数据方法常损害语义质量。本文指出现有技术忽视了提示输入中蕴含的潜在信号,且仅关注样本内偏好差异,忽略样本间差异。为此,提出多层级感知偏好学习(MAPL)框架:针对原始偏好数据中的每个响应,构造不同条件下的变体提示以学习样本内偏好差异;同时,将原始偏好对扩展为多指令偏好对,捕获样本间偏好差异。基于这两类数据,设计两个精细化训练目标函数,并可无缝集成至奖励建模与直接偏好优化范式。在多个基准上的实证评估验证了该框架的有效性。

原文摘要 · Abstract (English)

RLHF has emerged as a predominant approach for aligning artificial intelligence systems with human preferences, demonstrating exceptional and measurable efficacy in instruction following tasks; however, it exhibits insufficient compliance capabilities when confronted with complex multi-instruction tasks. Conventional approaches rely heavily on human annotation or more sophisticated large language models, thereby introducing substantial resource expenditure or potential bias concerns. Meanwhile, alternative synthetic methods that augment standard preference datasets often compromise the model's semantic quality. Our research identifies a critical oversight in existing techniques, which predominantly focus on comparing responses while neglecting valuable latent signals embedded within prompt inputs, and which only focus on preference disparities at the intra-sample level, while neglecting to account for the inter-sample level preference differentials that exist among preference data. To leverage these previously neglected indicators, we propose a novel Multi-level Aware Preference Learning (MAPL) framework, capable of enhancing multi-instruction capabilities. Specifically, for any given response in original preference data pairs, we construct varied prompts with a preference relation under different conditions, in order to learn intra-sample level preference disparities. Furthermore, for any given original preference pair, we synthesize multi-instruction preference pairs to capture preference discrepancies at the inter-sample level. Building on the two datasets constructed above, we consequently devise two sophisticated training objective functions. Subsequently, our framework integrates seamlessly into both Reward Modeling and Direct Preference Optimization paradigms. Through rigorous evaluation across multiple benchmarks, we empirically validate the efficacy of our framework.

RLHF指令遵循偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。