用离线世界模型优化音乐推荐,让情绪变好更安全可靠。
Affective Music Recommendation: A Rollout-Based World Model for Offline Preference Optimization
- 用因果Transformer建模听歌行为与情绪变化,模拟真实反馈。
- 离线微调后情绪预测提升,且避免了单一偏好导致的多样性丢失。
- 适合临床人群或伦理受限场景下的情感化推荐系统设计。
功能性音乐应用(如专注、助眠、情绪调节)的成功取决于听众的情绪状态,但在线实验存在伦理限制,尤其对神经认知障碍等临床人群。我们提出AMRS系统,部署于LUCID健康平台,服务临床用户(主要为老年认知障碍者)及普通用户,覆盖唤醒、专注、平静、睡眠四种模式。该系统基于滚动式世界模型:一个在日志数据上训练的因果Transformer,联合预测参与度、评分结果以及自报告的效价与唤醒度。该模型既用于离线策略训练的仿真,也作为部署前的压力测试工具。推荐策略以行为克隆初始化,再通过直接偏好优化(DPO)在离线环境中针对可配置的多目标效用函数进行微调。在严格冷启动条件下,世界模型能以可用精度预测行为与情绪信号;相较基线,DPO提升了预测的效价与唤醒度,同时保持相似的多样性分布,未出现贪婪优化带来的分布坍缩。本工作验证了一种在无法进行在线实验时,实现情感推荐的有效方法。
原文摘要 · Abstract (English)
Functional music applications, from consumer focus and sleep aids to clinical interventions, share a distinctive recommendation problem: success is defined by the listener's affective state, but online experimentation on emotion is ethically constrained, particularly for clinical populations who cannot reliably skip a song or report distress. We describe AMRS, the Affective Music Recommendation System deployed on LUCID's health-and-wellness platforms, which serve clinical users (primarily older adults with neurocognitive conditions) and consumer-wellness users across energize, focus, calm, and sleep modes. AMRS is built around a rollout-based world model: a causal transformer trained on logged listening data to jointly predict engagement, binary rating, and self-reported valence and arousal. The world model serves both as an in-silico simulator for offline policy training and as a stress-testing tool before deployment. A recommender policy initialized by behaviour cloning is fine-tuned offline with Direct Preference Optimization (DPO) against a configurable multi-objective utility function. Under a strict cold-start protocol, the world model predicts both behavioural and affective signals with usable fidelity; DPO improves predicted valence and arousal over the cloned baseline while maintaining a similar diversity profile and avoiding the distributional collapse produced by greedy optimization. We position the work as an early deployed validation of a methodology for affective recommendation when online experimentation is ethically untenable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。