arXiv:2508.07638cs.LG2025-08被引 1

用细粒度偏好数据精选提升大模型对齐效果

Data Selection for LLM Alignment Using Fine-Grained Preferences

  • 将对齐问题转化为数据选择,筛选冲突最严重的样本
  • 仅用30%数据就超越全量数据训练效果
  • 适合需要高效对齐且有细粒度标注的场景

大语言模型对齐旨在使模型行为符合人类偏好。随着多维度、细粒度偏好数据的获取日益可行,现有对齐方法通常仅处理单一偏好,难以应对此类数据集固有的冲突。本文提出一种以数据为中心的方法,通过直接优化细粒度偏好,引入偏好分歧(PD)量化跨维度偏好冲突。为避免复杂优化,将问题重构为数据选择,提出一种简单有效策略:选取具有最负PD值的数据子集进行训练。理论分析表明该策略具有损失边界最优性。在多种设置和数据集上的实证研究显示,该方法仅使用30%数据即可持续优于标准全数据对齐。本工作验证了基于细粒度偏好的大模型对齐高度可行。

原文摘要 · Abstract (English)

Large language models (LLMs) alignment aims to ensure that the behavior of LLMs meets human preferences. While collecting data from multiple fine-grained, aspect-specific preferences becomes more and more feasible, existing alignment methods typically work on a single preference and thus struggle with conflicts inherent in such aggregated datasets. As one early attempt, in this paper, we propose a data-centric approach to align LLMs through the effective use of fine-grained preferences. Specifically, we formulate the problem as a direct fine-grained preference optimization and introduce preference divergence (PD) that quantifies inter-aspect preference conflicts. Instead of directly tackling the consequent complicated optimization, we recast it as a data selection problem and propose a simple yet effective strategy, which identifies a subset of data corresponding to the most negative PD values, for efficient training. We theoretically analyze the loss-bound optimality of our selection strategy and conduct extensive empirical studies on varied settings and datasets to demonstrate that our practical selection method could achieve consistent improvement against standard full-data alignment, using even just 30% of the data. Our work shares a line that LLM alignment using fine-grained preferences is highly feasible.

大模型对齐偏好学习数据选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。