arXiv:2505.23557stat.MLcs.LG2025-05ICML

用偏好信息提升参数估计精度,可实现更快收敛速度。

Learning Parametric Distributions from Samples and Preferences

  • 基于样本与偏好构建新估计器,利用偏好信息优化参数学习
  • 估计误差可达 $\mathcal{O}(1/n)$,优于仅用样本的 $\Theta(1/\sqrt{n})$
  • 适用于高斯、拉普拉斯等分布,适合偏好反馈场景建模

近期语言建模进展凸显了偏好反馈在提升模型性能中的作用。本文研究在何种条件下,偏好反馈能改善连续参数分布类中的参数估计。在该框架中,学习者观察来自未知分布的样本对及其相对偏好,而偏好依赖于同一未知参数。我们证明,基于偏好的M-估计器的渐近方差优于仅使用样本的M-估计器,且确定性偏好进一步提升性能。利用确定性偏好带来的硬约束,我们提出一种估计器,其估计误差呈 $\mathcal{O}(1/n)$ 阶,显著优于仅用样本可达到的 $\Theta(1/\sqrt{n})$。随后,我们建立了匹配该加速率的下界,至多相差维度与问题相关常数。尽管分析假设较严格,但对基于对数概率奖励的高斯或拉普拉斯分布等典型情形成立。

原文摘要 · Abstract (English)

Recent advances in language modeling have underscored the role of preference feedback in enhancing model performance. This paper investigates the conditions under which preference feedback improves parameter estimation in classes of continuous parametric distributions. In our framework, the learner observes pairs of samples from an unknown distribution along with their relative preferences depending on the same unknown parameter. We show that preference-based M-estimators achieve a better asymptotic variance than sample-only M-estimators, further improved by deterministic preferences. Leveraging the hard constraints revealed by deterministic preferences, we propose an estimator achieving an estimation error scaling of $\mathcal{O}(1/n)$ -- a significant improvement over the $Θ(1/\sqrt{n})$ rate attainable with samples alone. Next, we establish a lower bound that matches this accelerated rate; up to dimension and problem-dependent constants. While the assumptions underpinning our analysis are restrictive, they are satisfied by notable cases such as Gaussian or Laplace distributions for preferences based on the log-probability reward.

参数估计偏好学习统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。