从多轮交互中学习个性化目标,无需直接奖励信号
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
- 基于隐式反馈构建多步奖励估计器,解决马尔可夫决策过程中的反馈难题
- 提出高效算法,在合成与真实用户预订数据上实现亚线性累积误差
- 适合研究大语言模型多轮交互、个性化推荐等场景的从业者
本文研究交互基础学习(IGL),一种在现实场景中面对未知机制生成的间接反馈而非显式数值奖励的学习范式。尽管已有IGL工作提供了具有理论保障的高效算法,但其结果局限于单步设置,难以适用于现代序列决策系统(如多轮大语言模型部署)。为此,我们提出一种计算高效的算法,实现了具有个性化反馈的上下文时期性马尔可夫决策过程(MDP)的亚线性遗憾保证。技术上,我们将Zhang等(2024a)的奖励估计器构造从单步扩展至多步设置,解决了在MDP下解码潜在奖励的独特挑战。基于该估计器,设计了逆间隙加权(IGW)策略优化算法。最后,通过在合成时期性MDP和真实用户预订数据集上的实验,验证了方法在从多轮交互中学习个性化目标的有效性。
原文摘要 · Abstract (English)
In this paper, we study Interaction-Grounded Learning (IGL) [Xie et al., 2021], a paradigm designed for realistic scenarios where the learner receives indirect feedback generated by an unknown mechanism, rather than explicit numerical rewards. While prior work on IGL provides efficient algorithms with provable guarantees, those results are confined to single-step settings, restricting their applicability to modern sequential decision-making systems such as multi-turn Large Language Model (LLM) deployments. To bridge this gap, we propose a computationally efficient algorithm that achieves a sublinear regret guarantee for contextual episodic Markov Decision Processes (MDPs) with personalized feedback. Technically, we extend the reward-estimator construction of Zhang et al. [2024a] from the single-step to the multi-step setting, addressing the unique challenges of decoding latent rewards under MDPs. Building on this estimator, we design an Inverse-Gap-Weighting (IGW) algorithm for policy optimization. Finally, we demonstrate the effectiveness of our method in learning personalized objectives from multi-turn interactions through experiments on both a synthetic episodic MDP and a real-world user booking dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。