arXiv:2503.08035cs.CL2025-03被引 1

让大模型按不同用户群体偏好生成回应,提升个性化效果。

Group Preference Alignment: Customized LLM Response Generation from In-Situ Conversations

  • 从真实对话日志中提取群体偏好差异,生成可解释的规则
  • 通过动态提示或微调合成数据,使模型适配不同群体需求
  • 在保持基准性能前提下,显著提升对群体偏好的响应契合度

大语言模型因通用训练范式难以满足不同用户群体的特殊需求,且缺乏对各群体个性化期望的研究。为此,本文提出群体感知个性化框架 Group Preference Alignment (GPA),识别用户群体在对话中的上下文相关偏好差异,并引导模型适配这些偏好。该方法包含两步:(1) 群体感知偏好提取,从真实对话日志中提取最大差异化的群体偏好,提炼为可解释的规则;(2) 定制化响应生成,通过两种方式实现:a) 上下文调优推理(GAP-CT),基于上下文动态调整提示指令以生成响应;b) 规则微调推理(GPA-FT),利用规则生成对比性合成数据,通过对齐方式微调特定群体模型。实验表明,该框架显著提升输出与用户偏好的契合度,优于基线方法,同时在标准基准测试上保持稳健表现。

原文摘要 · Abstract (English)

LLMs often fail to meet the specialized needs of distinct user groups due to their one-size-fits-all training paradigm \cite{lucy-etal-2024-one} and there is limited research on what personalization aspects each group expect. To address these limitations, we propose a group-aware personalization framework, Group Preference Alignment (GPA), that identifies context-specific variations in conversational preferences across user groups and then steers LLMs to address those preferences. Our approach consists of two steps: (1) Group-Aware Preference Extraction, where maximally divergent user-group preferences are extracted from real-world conversation logs and distilled into interpretable rubrics, and (2) Tailored Response Generation, which leverages these rubrics through two methods: a) Context-Tuned Inference (GAP-CT), that dynamically adjusts responses via context-dependent prompt instructions, and b) Rubric-Finetuning Inference (GPA-FT), which uses the rubrics to generate contrastive synthetic data for personalization of group-specific models via alignment. Experiments demonstrate that our framework significantly improves alignment of the output with respect to user preferences and outperforms baseline methods, while maintaining robust performance on standard benchmarks.

大模型个性化对话系统群体偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。