arXiv:2410.16027cs.CL2024-10NAACL被引 22

用社区偏好个性化大模型,让回答更符合特定群体口味。

ComPO: Community Preferences for Language Model Personalization

  • 根据社区标识调整模型输出分布,实现群体化偏好优化。
  • 在Reddit社区数据上,使用正确社区标签性能显著提升。
  • 适合需要适配特定用户群体的场景,如社群对话系统。

传统基于人类反馈训练语言模型的方法依赖于假设代表‘平均用户’的偏好,忽略了主观差异和细粒度变化。近期研究指出,将多样且常矛盾的人类反馈聚合用于微调模型,会导致生成结果偏向通用性,无法满足多数用户群体的偏好,因风格与规范被平均化。为此,我们借鉴推荐系统思想,提出ComPO方法,通过将偏好提供者的上下文信息融入模型输出的概率分布,实现语言模型的偏好个性化。聚焦群体层面而非个体,我们收集并发布了ComPRed——一个来自Reddit的问答数据集,包含社区级偏好,可在不引发隐私问题的情况下研究偏好多样性。实验表明,在偏好微调阶段引入社区标识(即子版块名称)能显著提升模型表现;而以随机子版块名替代则导致性能大幅下降,证明了该方法在适配社区偏好的有效性。

原文摘要 · Abstract (English)

Conventional algorithms for training language models (LMs) with human feedback rely on preferences that are assumed to account for an "average" user, disregarding subjectivity and finer-grained variations. Recent studies have raised concerns that aggregating such diverse and often contradictory human feedback to finetune models results in generic models that generate outputs not preferred by many user groups, as they tend to average out styles and norms. To address this issue, we draw inspiration from recommendation systems and propose ComPO, a method to personalize preference optimization in LMs by contextualizing the probability distribution of model outputs with the preference provider. Focusing on group-level preferences rather than individuals, we collect and release ComPRed, a question answering dataset with community-level preferences from Reddit. This dataset facilitates studying diversity in preferences without incurring privacy concerns associated with individual feedback. Our experiments reveal that conditioning language models on a community identifier (i.e., subreddit name) during preference tuning substantially enhances model performance. Conversely, replacing this context with random subreddit identifiers significantly diminishes performance, highlighting the effectiveness of our approach in tailoring responses to communities' preferences.

语言模型偏好优化社区建模个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。