构建首个多维奖励评估基准,提升智能推荐系统理解复杂意图能力
RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems

- 设计多维度奖励模型,覆盖指令遵循、事实一致等四类能力
- 包含超百万条数据,支持从语法合规到意图理解的全链条评估
- 适合研究智能推荐系统、强化学习奖励建模的学者与工程师
大型语言模型代理正推动推荐系统从简单的查询-物品匹配向深度个性化与交互式推荐演进。强化学习为优化这类代理在推荐任务中的表现提供了关键框架,但现有方法仍受限于单一维度的结果型奖励,仅关注最终用户行为,忽视了指令遵循、复杂意图理解等中间能力。尽管多维奖励设计至关重要,领域内尚无标准化评估基准。为此,我们提出 RecRM-Bench,目前规模最大、最全面的智能推荐系统评估基准。该基准包含超过100万条结构化数据,涵盖四大核心评估维度:指令遵循、事实一致性、查询-物品相关性,以及细粒度用户行为预测。通过支持从句法合规到复杂意图对齐和偏好建模的全流程评估,RecRM-Bench为训练复杂奖励模型提供了基础数据集。此外,我们提出一套系统化的多维奖励模型构建框架及混合奖励函数集成方法,为开发可靠且高性能的智能推荐系统奠定坚实基础。完整 RecRM-Bench 数据集已公开发布于 https://huggingface.co/datasets/wwzeng/RecRM-Bench。
原文摘要 · Abstract (English)
The integration of Large Language Model (LLM) agents is transforming recommender systems from simple query-item matching towards deeply personalized and interactive recommendations. Reinforcement Learning (RL) provides an essential framework for the optimization of these agents in recommendation tasks. However, current methodologies remain limited by a reliance on single dimensional outcome-based rewards that focus exclusively on final user interactions, overlooking critical intermediate capabilities, such as instruction following and complex intent understanding. Despite the necessity for designing multi-dimensional reward, the field lacks a standardized benchmark to facilitate this development. To bridge this gap, we introduce RecRM-Bench, the largest and most comprehensive benchmark to date for agentic recommender systems. It comprises over 1 million structured entries across four core evaluation dimensions: instruction following, factual consistency, query-item relevance, and fine-grained user behavior prediction. By supporting comprehensive assessment from syntactic compliance to complex intent grounding and preference modeling, RecRM-Bench provides a foundational dataset for training sophisticated reward models. Furthermore, we propose a systematic framework for the construction of multi-dimensional reward models and the integration of a hybrid reward function, establishing a robust foundation for developing reliable and highly capable agentic recommender systems. The complete RecRM-Bench dataset is publicly available at https://huggingface.co/datasets/wwzeng/RecRM-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。