arXiv:2601.19197cs.IRcs.AI2026-01

提出一套评估大模型推荐系统的人类中心框架,揭示传统指标忽略的体验短板。

HELM: A Human-Centered Evaluation Framework for LLM-Powered Recommender Systems

  • 从意图对齐、解释质量等五维度构建人本评估体系
  • 发现GPT-4解释力强但存在显著热门偏好(基尼系数0.73)
  • 适合关注用户体验与公平性的推荐系统研究者使用

大语言模型(LLM)融入推荐系统带来了自然语言理解、解释生成和对话交互等新能力,但现有评估方法仍聚焦传统准确率指标,难以捕捉影响真实用户体验的多维人本特征。本文提出 ramework{}(Human-centered Evaluation for LLM-powered Recommenders),一个涵盖意图对齐、解释质量、交互自然度、信任透明度及公平多样性五个维度的综合评估框架。通过在电影、图书、餐厅三个领域,对GPT-4、LLaMA-3.1和P5三种先进基于LLM的推荐系统进行测试,由12位领域专家在847个推荐场景下完成严格评估,结果表明: ramework{} 能揭示传统指标无法发现的关键质量问题。实验显示,尽管GPT-4在解释质量(4.21/5.0)和交互自然度(4.35/5.0)上表现优异,其热门偏好程度(基尼系数0.73)远高于传统协同过滤(0.58)。本文开源了该评估工具包,以推动推荐系统领域的人本评估实践。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) into recommendation systems has introduced unprecedented capabilities for natural language understanding, explanation generation, and conversational interactions. However, existing evaluation methodologies focus predominantly on traditional accuracy metrics, failing to capture the multifaceted human-centered qualities that determine the real-world user experience. We introduce \framework{} (\textbf{H}uman-centered \textbf{E}valuation for \textbf{L}LM-powered reco\textbf{M}menders), a comprehensive evaluation framework that systematically assesses LLM-powered recommender systems across five human-centered dimensions: \textit{Intent Alignment}, \textit{Explanation Quality}, \textit{Interaction Naturalness}, \textit{Trust \& Transparency}, and \textit{Fairness \& Diversity}. Through extensive experiments involving three state-of-the-art LLM-based recommenders (GPT-4, LLaMA-3.1, and P5) across three domains (movies, books, and restaurants), and rigorous evaluation by 12 domain experts using 847 recommendation scenarios, we demonstrate that \framework{} reveals critical quality dimensions invisible to traditional metrics. Our results show that while GPT-4 achieves superior explanation quality (4.21/5.0) and interaction naturalness (4.35/5.0), it exhibits a significant popularity bias (Gini coefficient 0.73) compared to traditional collaborative filtering (0.58). We release \framework{} as an open-source toolkit to advance human-centered evaluation practices in the recommender systems community.

推荐系统大模型评估人本设计公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。