arXiv:2501.06980cs.LGcs.AI2025-01被引 7

用大模型实时理解用户偏好,优化强化学习干预策略

Combining LLM decision and RL action selection to improve RL policy for adaptive interventions

  • 大模型分析用户文本偏好,动态过滤强化学习动作选择
  • 在模拟环境中验证,策略个性化程度显著提升
  • 适合医疗自适应干预、人机交互等需要实时响应的场景

强化学习在个性化健康自适应干预中日益重要。受大型语言模型(LLM)成功启发,我们探索利用LLM实时更新强化学习策略,以加速个性化进程。通过文本形式的用户偏好(如个人喜好、健康状态、约束条件等)实时影响动作选择,实现即时个性化调整。提出一种混合方法:将LLM输出作为典型强化学习动作选择的过滤器。研究了多种提示策略与动作选择机制。在生成文本偏好并建模行为动态约束的模拟环境中评估,结果表明该方法能有效融合文本偏好,同时改进强化学习策略,提升自适应干预的个性化水平。

原文摘要 · Abstract (English)

Reinforcement learning (RL) is increasingly being used in the healthcare domain, particularly for the development of personalized health adaptive interventions. Inspired by the success of Large Language Models (LLMs), we are interested in using LLMs to update the RL policy in real time, with the goal of accelerating personalization. We use the text-based user preference to influence the action selection on the fly, in order to immediately incorporate the user preference. We use the term "user preference" as a broad term to refer to a user personal preference, constraint, health status, or a statement expressing like or dislike, etc. Our novel approach is a hybrid method that combines the LLM response and the RL action selection to improve the RL policy. Given an LLM prompt that incorporates the user preference, the LLM acts as a filter in the typical RL action selection. We investigate different prompting strategies and action selection strategies. To evaluate our approach, we implement a simulation environment that generates the text-based user preferences and models the constraints that impact behavioral dynamics. We show that our approach is able to take into account the text-based user preferences, while improving the RL policy, thus improving personalization in adaptive intervention.

强化学习大模型个性化干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。