arXiv:2609.01627cs.IRcs.AI2026-09

用大模型评估推荐系统解释效果,发现其可靠但需谨慎使用。

The Utility of LLMs in Recommender Systems Explanation Evaluation

论文配图:The Utility of LLMs in Recommender Systems Explanation Evaluation
图 1 · 摘自论文原文
  • 用不同提示词生成18种解释原型,让14个大模型评分
  • 大模型评分与人类相似但绝对一致性差,大小和设置影响结果
  • 建议简洁提示、选大模型、测试评价框架、检查事实准确性

解释在构建可信推荐系统中至关重要,但选择合适的解释方法仍具挑战。现有方法常生成抽象输出,需额外处理才能用户友好,选项繁多难以逐一评测。人工评测不切实际,自动化指标或仅评估输出抽象性,或需真实标签(通常不可得)。近期研究显示大语言模型可充当解释评估的‘裁判’,但其可靠性尚未充分验证。本文研究大模型在为特定应用场景选择有效解释方法中的作用。我们生成18种不同解释原型,由14个不同规模的大模型在两种温度设置下评分,并与用户研究的人类评分对比。结果显示,大模型评分模式接近人类,与人类评分有中等排名相关性,但绝对评分一致性低,且随模型大小和评价结构差异显著。据此提出四条实用建议:保持解释生成提示简洁,优先使用更大模型进行评估,预先测试评价结构,审计解释的事实准确性,因人类与大模型均难识别非事实内容。

原文摘要 · Abstract (English)

Explanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method presents challenges. Many explanation methods exist, but little guidance exists on which is best for which setting. Existing explanation generation methods often produce abstract outputs that require further formatting to become user-friendly, with a seemingly endless pool of options. Running user-based evaluations of all possible options is usually unfeasible, while automated evaluation metrics often either assess only the explainer's abstract output or require comparison with a ground truth, which is generally unavailable. Recent studies have shown that large language models (LLMs) can serve as ``judges'' for explanation evaluation, but their reliability has not yet been thoroughly explored. This paper studies the utility of LLMs in selecting an effective explanation method for a given application. We first explore their ability to generate explanation prototypes given varying information about the RS and the user. Specifically, we generate 18 distinct explanation prototypes, which are subsequently evaluated by 14 LLMs of varying sizes across two temperature settings. We compare these against human ratings derived from a user study. Our results show that while LLMs exhibit human-like rating patterns and achieve moderate rank correlation with human raters, their absolute rating agreement is low and varies substantially by model size and evaluation construct. We derive four practical recommendations: keep explanation-generation prompts concise, prefer larger models for evaluation, pre-test evaluation constructs, and audit explanations for factual accuracy, as neither humans nor LLMs reliably detect non-factual content.

推荐系统大模型评估解释生成可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。