主动推断个性偏好能让小模型更准确地匹配用户需求。
Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?
- 用名人真实偏好构建合成数据,主动推断个性描述
- 使用1-8B模型测试,主动前缀显著提升泛化能力
- 减少对性别、年龄等属性的系统性偏差,适合个性化应用
语言模型对个性化偏好对齐的一个主要挑战是信息不足——用户未明确表达偏好。当前主流方法通过添加前置内容(如历史对话)来引导偏好分布,多为被动建模已有偏好对。本文探讨模型是否需主动推断偏好描述,构建基于知名人物公开偏好的合成个性化对齐数据集,测试1-8B规模模型在推断与对齐方面的表现。结果表明,高质量的主动前缀能带来更好的泛化效果、更符合上下文的输出,以及对不同受保护属性(如性别、年龄)更少的系统性偏差。所有结果均表明,主动对齐可为个性化对齐提供更可控、高效的新路径。
原文摘要 · Abstract (English)
A prominent issue in aligning language models (LMs) to personalized preferences is underspecification -- the lack of information from users about their preferences. A popular trend of injecting such specification is adding a prefix (e.g. prior relevant conversations) to the current user's conversation to steer preference distribution. Most methods passively model personal preferences with prior example preferences pairs. We ask whether models benefit from actively inferring preference descriptions, and address this question by creating a synthetic personalized alignment dataset based on famous people with known public preferences. We then test how effective finetuned 1-8B size models are at inferring and aligning to personal preferences. Results show that higher-quality active prefixes lead to better generalization, more contextually faithful models, and less systematic biases across different protected attributes. All our results suggest active alignment can lead to a more controllable and efficient path for personalized alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。