arXiv:2501.18282cs.LG2025-01ICML被引 2

利用稀疏性降低偏好学习样本需求,理论证明可大幅减少标注数据量。

Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective

  • 基于稀疏奖励模型,将误差率从Θ(d/n)优化至Θ(k/n log(d/k))
  • 在合成数据和大模型对齐数据上,稀疏方法显著提升精度并减少样本量
  • 适用于高维特征下标注成本高的偏好学习场景

本文研究偏好学习的样本效率问题,即基于对比判断建模与预测人类选择。经典估计理论中最小最大最优误差率为Θ(d/n),要求样本数n随特征维度d线性增长。然而,高维特征空间与人工标注数据的高昂成本制约了传统方法效率。为此,本文利用偏好模型中的稀疏性,建立了紧致的误差率。在奖励函数参数为k-稀疏的随机效用模型下,最小最大最优率可降至Θ(k/n log(d/k))。进一步分析了ℓ₁正则化估计器,在格拉姆矩阵满足温和假设下达到近最优率。在合成数据与大模型对齐数据上的实验验证了理论结果,表明稀疏感知方法能显著降低样本复杂度并提升预测准确率。

原文摘要 · Abstract (English)

This paper considers the sample-efficiency of preference learning, which models and predicts human choices based on comparative judgments. The minimax optimal estimation error rate $Θ(d/n)$ in classical estimation theory requires that the number of samples $n$ scales linearly with the dimensionality of the feature space $d$. However, the high dimensionality of the feature space and the high cost of collecting human-annotated data challenge the efficiency of traditional estimation methods. To remedy this, we leverage sparsity in the preference model and establish sharp error rates. We show that under the sparse random utility model, where the parameter of the reward function is $k$-sparse, the minimax optimal rate can be reduced to $Θ(k/n \log(d/k))$. Furthermore, we analyze the $\ell_{1}$-regularized estimator and show that it achieves near-optimal rate under mild assumptions on the Gram matrix. Experiments on synthetic data and LLM alignment data validate our theoretical findings, showing that sparsity-aware methods significantly reduce sample complexity and improve prediction accuracy.

偏好学习稀疏性样本效率理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。