用真实用户数据评估大模型健康教练,发现统一多工具策略伤害特定群体。
Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users
- 分解决策头(工具/风格)实现离线评估,识别策略优劣
- 轻量模拟器加早期信息增益奖励,提升特质识别速度与目标达成率
- 建议先评估再个性化,按子群体制定指标避免平均值掩盖问题
我们研究了一个在网页上部署、具备工具增强功能的大型语言模型健康教练系统,基于7名真实用户的试点数据(280个评分回合)。通过因子化决策头(工具/风格)进行离线策略评估(OPE),发现采用统一重工具策略虽提升整体平均价值,却对低健康素养但高自我效能的用户群体造成负面影响。一个包含隐式原型的轻量级模拟器进一步显示,加入微小的早期信息增益奖励可稳定缩短特质识别时间,并提升目标成功率与pass@3指标。这些初步结果表明,应采取‘评估先行’的个性化路径:冻结生成器,基于类型化奖励(客观工具结果与满意度)学习子群感知的决策头,并始终报告各原型的指标,以揭示被平均值掩盖的子群伤害。
原文摘要 · Abstract (English)
We study a web-deployed, tool-augmented LLM health coach with real users. In a pilot with seven users (280 rated turns), offline policy evaluation (OPE) over factorized decision heads (Tool/Style) shows that a uniform heavy-tool policy raises average value on logs but harms specific subgroups, most notably low-health-literacy/high-self-efficacy users. A lightweight simulator with hidden archetypes further shows that adding a small early information-gain bonus reliably shortens trait identification and improves goal success and pass@3. Together, these early findings indicate an evaluation-first path to personalization: freeze the generator, learn subgroup-aware decision heads on typed rewards (objective tool outcomes and satisfaction), and always report per-archetype metrics to surface subgroup harms that averages obscure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。