arXiv:2505.11861cs.AIcs.CL2025-05被引 4

构建个性化社会公平偏好数据集,提升大模型对多元群体的适配能力

Fair-PP: A Synthetic Dataset for Aligning LLM with Personalized Preferences of Social Equity

  • 基于真实社会调查生成28个群体、98个议题的合成偏好数据
  • 用角色扮演生成23万条偏好记录,实现跨区域模型偏好多维分析
  • 提出可重加权的对齐方法,精准匹配目标人群偏好

人类偏好在大语言模型优化中至关重要,但人工收集偏好成本高,且现有数据集普遍忽视个性化与偏好的关联。为此,我们提出 Fair-PP,一个面向社会公平的个性化偏好合成数据集,源自真实社会调查数据,涵盖28个社会群体、98个公平议题及5个个人偏好维度。利用 GPT-4o-mini,我们基于七种代表性人物设定进行角色扮演,生成共计238,623条偏好记录。通过 Fair-PP,我们贡献:(i) 一种自动生成偏好数据的框架,以及更细粒度的个性化偏好数据集;(ii) 对主流 LLM 在五大全球区域中的个性化偏好定位分析;(iii) 一种样本重加权方法,实现针对目标人格的偏好对齐,同时最大化与其他人格的差异。实验证明该方法优于基线。

原文摘要 · Abstract (English)

Human preference plays a crucial role in the refinement of large language models (LLMs). However, collecting human preference feedback is costly and most existing datasets neglect the correlation between personalization and preferences. To address this issue, we introduce Fair-PP, a synthetic dataset of personalized preferences targeting social equity, derived from real-world social survey data, which includes 28 social groups, 98 equity topics, and 5 personal preference dimensions. Leveraging GPT-4o-mini, we engage in role-playing based on seven representative persona portrayals guided by existing social survey data, yielding a total of 238,623 preference records. Through Fair-PP, we also contribute (i) An automated framework for generating preference data, along with a more fine-grained dataset of personalized preferences; (ii) analysis of the positioning of the existing mainstream LLMs across five major global regions within the personalized preference space; and (iii) a sample reweighting method for personalized preference alignment, enabling alignment with a target persona while maximizing the divergence from other personas. Empirical experiments show our method outperforms the baselines.

大模型对齐社会公平个性化偏好合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。