arXiv:2511.02966cs.LG2025-11NeurIPS被引 2

用少量对比提问,快速匹配用户偏好生成内容。

Inference-Time Personalized Alignment with a Few User Preference Queries

  • 通过少量成对比较获取用户偏好,动态筛选最优生成结果。
  • 在文本与图像生成任务中,仅需3~5次提问即实现精准对齐。
  • 适合需要快速个性化响应的交互式生成场景。

我们研究如何使生成模型的输出符合用户偏好。现有方法要么需要大量偏好查询,要么要求用户明确输入偏好描述。本文提出一种新的推理阶段个性化对齐方法 UserAlign,仅通过少量成对响应比较即可捕捉用户偏好。UserAlign 基于逻辑置信区间中的最优臂识别理论,从模型预生成的固定响应池中选择最符合用户偏好的结果。核心思想是假设用户反馈一致且无噪声,并将其融入理论框架以加速最佳响应识别。在多个任务上的实验表明,该方法在个性化文本与图像生成中均表现优异。

原文摘要 · Abstract (English)

We study the problem of aligning a generative model's response with a user's preferences. Recent works have proposed several different formulations for personalized alignment; however, they either require a large amount of user preference queries or require that the preference be explicitly specified as a text input. In this paper, we propose a novel inference-time personalized alignment method, UserAlign, that elicits the user's preferences with a few queries as pairwise response comparisons. In particular, UserAlign builds on the theoretical framework of best-arm identification in logistic bandits and selects a personalized response from a fixed pool of the model's generated responses. The key idea is to consider the user's feedback consistent and noise-free, and incorporate it into the theoretical framework to identify the best response quickly. Experimental results across several tasks, involving personalized text and image generation, showcase the effectiveness of UserAlign in achieving personalized alignment.

个性化生成偏好对齐少样本交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。