用大模型生成用户交互数据,辅助训练健康行为干预的强化学习模型。
Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
- 直接用大模型生成模拟用户交互数据,无需真实数据即可启动训练。
- 生成数据在效果上达到人工标注水平,且可替代真人评估样本。
- 不同提示策略效果差异大,需根据具体研究和模型调整使用。
个性化数字健康应用是提升用户参与度和干预效果的重要方向,尤其在动态适应用户状态(如动机、知识、需求)时更为有效。然而,此类系统的设计涉及大量决策,其效果难以从文献预测,实际验证又成本高昂。本文探索大语言模型(LLMs)是否可直接用于生成可用于训练强化学习模型的用户交互样本。基于四个大型行为改变研究的真实用户数据作为对照,我们发现:在缺乏真实数据时,LLM生成的样本仍具实用性;与人类标注者的样本对比,其表现相当。进一步分析显示,不同提示策略(包括短/长提示、思维链、少样本提示)的效果因研究任务和具体模型而异,仅提示语改写就可能带来显著差异。研究为如何在实践中有效利用LLM生成数据提供了具体建议。
原文摘要 · Abstract (English)
Personalizing digital applications for health behavior change is a promising route to making them more engaging and effective. This especially holds for approaches that adapt to users and their specific states (e.g., motivation, knowledge, wants) over time. However, developing such approaches requires making many design choices, whose effectiveness is difficult to predict from literature and costly to evaluate in practice. In this work, we explore whether large language models (LLMs) can be used out-of-the-box to generate samples of user interactions that provide useful information for training reinforcement learning models for digital behavior change settings. Using real user data from four large behavior change studies as comparison, we show that LLM-generated samples can be useful in the absence of real data. Comparisons to the samples provided by human raters further show that LLM-generated samples reach the performance of human raters. Additional analyses of different prompting strategies including shorter and longer prompt variants, chain-of-thought prompting, and few-shot prompting show that the relative effectiveness of different strategies depends on both the study and the LLM with also relatively large differences between prompt paraphrases alone. We provide recommendations for how LLM-generated samples can be useful in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。