arXiv:2501.11549cs.CL2025-01ACL被引 22

通过推理用户人格提升大模型个性化回应能力

Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas

  • 从偏好数据中推断用户人格,揭示选择原因
  • 结合人格信息训练后,模型个性化能力显著提升
  • 尤其擅长服务小众偏好用户,超越传统对齐方法

大语言模型通过学习用户对两个回复的偏好来对齐指令,但此类数据未说明用户偏好或拒绝的原因,导致模型无法针对不同用户需求定制回复。为此,本文提出双阶段方法:首先通过溯因推理(Persona Inference, PI)推断出支持被选或被拒回复的用户人格;其次通过人格定制(Persona Tailoring, PT)训练模型根据这些人格生成适配回复。实验表明:1)大模型能准确推断出解释用户偏好的人格特征;2)在偏好数据中加入PI生成的人格信息后,模型个性化能力增强,并可泛化至用户自定义人格;3)被拒回复对应的人格构成更具挑战性的评估场景,显示PT在服务非典型偏好用户方面优于传统对齐方法。本文主张以溯因视角理解偏好,不仅要问‘哪个更好’,更要追问‘何时、为何、为谁更好’。

原文摘要 · Abstract (English)

LLMs are aligned to follow input instructions by learning which of two responses users prefer for a prompt. However, such preference data do not convey why users prefer responses that are chosen or rejected, so LLMs trained on these datasets cannot tailor responses to varied user needs. To surface these parameters of personalization, we apply abductive reasoning to preference data, inferring needs and interests of users, i.e., personas, that may prefer either response. We test this idea in two steps: Persona Inference (PI), abductively inferring personas of users who prefer chosen or rejected outputs, and Persona Tailoring (PT), training models to tailor outputs to personas from PI. We show: 1) LLMs infer personas accurately explaining why different users may prefer both chosen or rejected outputs; 2) Training on preference data augmented with PI personas via PT boosts personalization and generalizes to supporting user-written personas; and 3) Rejected response personas form harder personalization evaluations, showing PT better aids users with uncommon preferences versus typical alignment methods. We argue for an abductive view of preferences for personalization, asking not only which response is better but when, why, and for whom.

个性化人格建模偏好学习溯因推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。