arXiv:2604.02527cs.LGcs.AI2026-04

LLM生成偏好数据能有效降低推荐系统早期损失,但噪声超过30%就失效。

Jump Start or False Start? A Theoretical and Empirical Evaluation of LLM-initialized Bandits

  • 用LLM生成偏好数据初始化上下文老虎机,减少初期试错
  • 当数据噪声超30%时,预热效果减弱;超50%时反而比冷启动差
  • 首次理论证明了LLM预热在何种对齐条件下更优,适合推荐系统研究者

大型语言模型(LLM)可生成用户偏好数据用于初始化上下文老虎机(CBLI),显著降低早期遗憾。然而,该方法依赖于生成选择与真实偏好的一致性。本文系统评估在随机噪声和标签翻转干扰下LLM生成偏好的表现:在对齐领域,预热在30%污染内仍有效,40%左右优势消失,50%以上性能反降;存在系统性偏差时,即使无噪声,其后悔值也高于冷启动。我们提出理论分析,分解随机标签噪声与系统性偏差对先验误差的影响,推导出在特定条件下基于LLM的预热优于冷启动的充分条件。在多个联合分析数据集与不同LLM上验证,估计的对齐度可准确预测预热效果的提升或恶化。

原文摘要 · Abstract (English)

The recent advancement of Large Language Models (LLMs) offers new opportunities to generate user preference data to warm-start bandits. Recent studies on contextual bandits with LLM initialization (CBLI) have shown that these synthetic priors can significantly lower early regret. However, these findings assume that LLM-generated choices are reasonably aligned with actual user preferences. In this paper, we systematically examine how LLM-generated preferences perform when random and label-flipping noise is injected into the synthetic training data. For aligned domains, we find that warm-starting remains effective up to 30% corruption, loses its advantage around 40%, and degrades performance beyond 50%. When there is systematic misalignment, even without added noise, LLM-generated priors can lead to higher regret than a cold-start bandit. To explain these behaviors, we develop a theoretical analysis that decomposes the effect of random label noise and systematic misalignment on the prior error driving the bandit's regret, and derive a sufficient condition under which LLM-based warm starts are provably better than a cold-start bandit. We validate these results across multiple conjoint datasets and LLMs, showing that estimated alignment reliably tracks when warm-starting improves or degrades recommendation quality.

推荐系统上下文老虎机大模型应用鲁棒性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。