arXiv:2512.19057cs.LGcs.IT2025-12

用最优实验设计减少人类反馈量,高效个性化生成模型

Efficient Personalization of Generative Models via Optimal Experimental Design

  • 通过最优实验设计选择最有效的人类偏好查询
  • 仅需较少反馈即可准确建模用户偏好,显著降低数据需求
  • 适合需要低成本个性化生成模型的场景

从人类反馈中进行偏好学习可将生成模型与终端用户需求对齐。然而,获取人类反馈成本高、耗时长,亟需数据高效的查询选择方法。本文提出一种新方法,利用最优实验设计来选择最具信息量的偏好查询,从而高效揭示建模用户偏好的潜在奖励函数。我们将偏好查询选择问题建模为最大化对底层偏好模型的信息量,证明该问题具有凸优化形式,并提出统计与计算高效算法 ED-PBRL,具备理论保证,可高效构造图像或文本等结构化查询。在个性化文生图模型以适应用户特定风格的实验中,结果表明其所需偏好查询数量显著少于随机选择。

原文摘要 · Abstract (English)

Preference learning from human feedback has the ability to align generative models with the needs of end-users. Human feedback is costly and time-consuming to obtain, which creates demand for data-efficient query selection methods. This work presents a novel approach that leverages optimal experimental design to ask humans the most informative preference queries, from which we can elucidate the latent reward function modeling user preferences efficiently. We formulate the problem of preference query selection as the one that maximizes the information about the underlying latent preference model. We show that this problem has a convex optimization formulation, and introduce a statistically and computationally efficient algorithm ED-PBRL that is supported by theoretical guarantees and can efficiently construct structured queries such as images or text. We empirically present the proposed framework by personalizing a text-to-image generative model to user-specific styles, showing that it requires less preference queries compared to random query selection.

生成模型偏好学习实验设计个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。