arXiv:2604.09553cs.IRcs.AI2026-04

构建首个面向大模型的序列推荐综合评测基准,解决评估不公问题。

SRBench: A Comprehensive Benchmark for Sequential Recommendation with Large Language Models

论文配图:SRBench: A Comprehensive Benchmark for Sequential Recommendation with Large Language Models
图 1 · 摘自论文原文
  • 设计多维评估框架,覆盖准确率、公平性、稳定性与效率。
  • 统一提示输入范式,实现神经网络与大模型的公平对比。
  • 创新提示-提取耦合机制,精准从大模型输出中抓取答案。

大语言模型(LLM)在序列推荐(SR)中的应用日益受到关注,但现有基准存在三大缺陷:1)过度关注准确率,忽视公平性等实际需求;2)数据集无法发挥LLM潜力,导致神经网络模型与LLM模型对比不公平;3)缺乏可靠机制从非结构化输出中提取任务特定答案。为此,我们提出SRBench,一个包含三项核心设计的综合性评测基准:1)多维度评估框架,涵盖准确率、公平性、稳定性和效率,贴合真实场景需求;2)通过提示工程实现统一输入范式,提升LLM-SR性能并保障模型间公平比较;3)提出提示-提取耦合机制,利用提示强制输出格式,并通过数值导向提取器精准捕获答案。我们使用SRBench评估了13个主流模型,发现LLM-SR模型过度关注物品流行度,缺乏对物品质量的深层理解。SRBench为序列推荐模型提供公平且全面的评估能力,支撑未来研究与应用。

原文摘要 · Abstract (English)

LLM development has aroused great interest in Sequential Recommendation (SR) applications. However, comprehensive evaluation of SR models remains lacking due to the limitations of the existing benchmarks: 1) an overemphasis on accuracy, ignoring other real-world demands (e.g., fairness); 2) existing datasets fail to unleash LLMs' potential, leading to unfair comparison between Neural-Network-based SR (NN-SR) models and LLM-based SR (LLM-SR) models; and 3) no reliable mechanism for extracting task-specific answers from unstructured LLM outputs. To address these limitations, we propose SRBench, a comprehensive SR benchmark with three core designs: 1) a multi-dimensional framework covering accuracy, fairness, stability and efficiency, aligned with practical demands; 2) a unified input paradigm via prompt engineering to boost LLM-SR performance and enable fair comparisons between models; 3) a novel prompt-extractor-coupled extraction mechanism, which captures answers from LLM outputs through prompt-enforced output formatting and a numeric-oriented extractor. We have used SRBench to evaluate 13 mainstream models and discovered some meaningful insights (e.g., LLM-SR models overfocus on item popularity but lack deep understanding of item quality). Concisely, SRBench enables fair and comprehensive assessments for SR models, underpinning future research and practical application.

序列推荐大模型评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。