用强化学习优化饮食推荐,平衡健康与偏好。
MORL-A2C: Multi-Objective Reinforcement Learning Reranker for Optimizing Healthiness in MOPI-HFRS

- 基于强化学习构建多目标序列重排序模型,动态调整推荐策略。
- 健康评分提升23.5个百分点,仅轻微降低推荐精度。
- 适合关注健康饮食的个性化推荐系统研究者。
不健康的饮食行为仍是美国持续存在的公共健康问题,部分原因在于推荐系统过度侧重用户偏好而忽视营养健康。本文扩展了多目标个性化可解释健康感知食物推荐系统(MOPI-HFRS),提出MORL-A2C,一种针对健康-偏好权衡的序列决策方法。该方法利用冻结的图神经网络嵌入,将推荐建模为K步重排序问题,采用优势演员-评论家算法并设计标量化的相关性/健康奖励。策略通过行为克隆从冻结嵌入的点积排序器中冷启动。我们还发现并修正了MOPI-HFRS评估流程中的关键漏洞,导致基线性能被低估;所有结果均基于修正后的基线。在宏营养素基准上,MORL-A2C使推荐质量略有下降(Recall@20: 25.64% → 23.61%,NDCG@20: 23.52% → 20.64%),但健康对齐度显著提升(H-Score@20: 46.05% → 69.57%),全营养素基准也呈现一致趋势。结果表明,基于策略的序列优化能有效处理多目标食品推荐中的健康-偏好权衡。
原文摘要 · Abstract (English)
Unhealthy dietary behavior continues to be a persistent public health issue in the United States, exacerbated by recommendation systems that prioritize user preference without considering nutritional health. The Multi-Objective Personalized Interpretable Health-aware Food Recommendation System (MOPI-HFRS), from which this work extends, addresses this by jointly optimizing preference, health, and diversity through Pareto-based optimization. However, this approach relies on static, per-step tradeoff solutions that fail to capture the sequential nature of dietary decision-making. We introduce MORL-A2C, a sequential decision-making extension to MOPI-HFRS targeting the health-preference axis. Leveraging frozen GNN embeddings, MORL-A2C formulates recommendation as a K-step reranking problem using an Advantage Actor-Critic algorithm with a scalarized relevance/health reward. The policy is warm-started via behavior cloning against a dot-product ranker derived from frozen embeddings. We also identify and correct a non-trivial bug in the MOPI-HFRS evaluation pipeline that understated baseline performance; all results are reported against the corrected baseline. On the macro-nutrient benchmark, MORL-A2C achieves a modest reduction in ranking quality (Recall@20: 25.64% to 23.61%, NDCG@20: 23.52% to 20.64%) in exchange for a substantial improvement in health alignment (H-Score@20: 46.05% to 69.57%), with consistent trends on the full-nutrient benchmark. These findings validate that policy-driven sequential optimization can effectively navigate the health-preference trade-off in multi-objective food recommendation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。