用户多样性是个性化大模型高效适配的关键
Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity

- 提出用户多样性是实现最优个性化学习的必要条件
- 在满足多样性条件下,贪婪算法可达最优效率
- 揭示用户偏好差异对模型可识别性的决定性作用
个性化对齐旨在适应异质用户偏好,但其统计效率的理论条件尚未明确。本文刻画了实现O(1)在线后悔率和log(1/epsilon)离线样本复杂度的条件:用户特定头部必须覆盖可能改变最优响应的潜在奖励方向。证明该条件既必要又充分。满足时,简单贪心算法达到基准效率;不满足时,自然可接受类中的所有学习者至少面临对数级后悔。结果表明,用户多样性是个性化可识别性的根本驱动力。
原文摘要 · Abstract (English)
Personalized alignment aims to adapt large language models to heterogeneous user preferences, yet the precise theoretical conditions for its statistical efficiency have not been formally established. This paper characterizes the conditions under which personalized alignment achieves O(1) online regret and log(1/epsilon) offline sample complexity. We show that these optimal rates depend on a specific user-diversity condition: the population of user-specific heads must span the latent reward directions that can alter the optimal response. We prove that this condition is both necessary and sufficient. When it holds, simple greedy algorithms achieve benchmark efficiency; when it fails, every learner in a natural admissible class incurs at least logarithmic regret. Our results identify user diversity as the fundamental driver of personalized identifiability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。