基于上下文动态调整,实时学习人类偏好并评估模型不确定性。
Contextual Online Uncertainty-Aware Preference Learning for Human Feedback
- 分两阶段决策:先探索后利用,结合ε-贪婪与自适应策略
- 在依赖性偏好数据下实现最优后悔率与估计量渐近正态性
- 适合需实时反馈、关注模型置信度的AI对齐场景
从人类反馈中进行强化学习(RLHF)已成为对齐大模型与人类偏好的关键范式。本文提出一种新型统计框架,基于动态上下文信息,同时实现在线决策与最优模型的统计推断。方法上采用两阶段算法:先ε-贪婪探索,再进入利用阶段;理论上通过定制反浓度不等式与矩阵鞅集中技术,推导出依赖样本下的统一估计速率及估计量渐近正态性。大量模拟结果表明,该方法优于现有先进策略。我们将该框架应用于大规模多任务语言理解数据集(MMLU)中对大语言模型医学解剖知识排序的人类偏好数据,揭示了不同模型在该任务上的表现差异。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm in artificial intelligence to align large models with human preferences. In this paper, we propose a novel statistical framework to simultaneously conduct the online decision-making and statistical inference on the optimal model using human preference data based on dynamic contextual information. Our approach introduces an efficient decision strategy that achieves both the optimal regret bound and the asymptotic distribution of the estimators. A key challenge in RLHF is handling the dependent online human preference outcomes with dynamic contexts. To address this, in the methodological aspect, we propose a two-stage algorithm starting with $ε$-greedy followed by exploitations; in the theoretical aspect, we tailor anti-concentration inequalities and matrix martingale concentration techniques to derive the uniform estimation rate and asymptotic normality of the estimators using dependent samples from both stages. Extensive simulation results demonstrate that our method outperforms state-of-the-art strategies. We apply the proposed framework to analyze the human preference data for ranking large language models on the Massive Multitask Language Understanding dataset, yielding insightful results on the performance of different large language models for medical anatomy knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。