测试大模型在重复推荐中从经验中学习的能力,发现现有模型难提升跨轮次表现。
Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
- 构建模拟用户互动的可控测试平台,聚焦跨轮次经验积累
- 模型能单轮内适应,但跨轮次策略优化能力不足
- 适合研究持续学习与个性化推荐的学者参考
为应对不断变化的真实环境,智能体需在知识不全时通过经验调整策略。然而当前对基于大模型的智能体评估大多忽略这一能力。本文强调不仅要在单个任务中应对不确定性(信息随对话逐步揭示),更要能在相似任务间通过积累经验、推断共有的潜在结构来改进适应能力。重复产品推荐提供了一个理想的实验场景:每轮中推荐系统需通过提问挖掘未知用户偏好;多轮后应根据观察到的用户与商品分布调整提问策略。为此,我们构建了体验式学习与主动探索基准(BELA),融合亚马逊真实商品目录、多样化的合成用户角色(捕捉异质潜在偏好)以及基于大模型的用户模拟框架,以模拟偏好揭示型交互。不同于完全复现真实行为,BELA旨在测试智能体是否能利用跨轮次的一致潜在偏好进行优化。基准测试显示,当前模型虽能在单轮内适应,却难以在多轮间实现策略改进,凸显了提升体验式学习能力的迫切需求。
原文摘要 · Abstract (English)
To navigate ever-shifting real-world environments, agents must grapple with incomplete knowledge and adapt their strategies through experience. However, current evaluations of LLM-based agents largely overlook this capability. Crucially, we stress not just the ability to contend with uncertainty within a task (episode), as episode-specific information is progressively revealed across turns, but also the refinement of such adaptive ability across similar episodes, as agents accumulate experiences and infer their shared latent structure. Repeated product recommendation offers a natural setting to isolate this need for experiential learning: within each interaction, an intelligent recommender must elicit unknown customer preferences through questions; across multiple interactions, it should tailor its questioning strategy to the observed customer and product distributions. We instantiate the Benchmark for Experiential Learning and Active exploration (BELA) by combining (1) a rich catalog of real-world products from Amazon, (2) a diverse collection of synthetic customer personas aimed to capture heterogeneous latent preferences, and (3) an LLM-based customer simulator framework that emulate preference-revealing interactions. Rather than aiming to faithfully replicate real consumer behavior, BELA provides a controlled and scalable testbed for whether agents can exploit consistent latent preferences across episodes. Benchmarking current models reveals that they can learn across turns, but struggle to improve across episodes. This underscore the need for frontier models to advance in experiential learning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。