arXiv:2507.13705cs.CLcs.IR2025-07中稿 · the Nineteenth ACM…被引 8

LLM推荐看似合理,解释却常自相矛盾。

Consistent Explainers or Unreliable Narrators? Understanding LLM-generated Group Recommendations

  • LLM推荐倾向加性功利聚合,但解释却模糊提及平均评分
  • 解释中频繁引入未定义的相似度、多样性等额外标准
  • 群体结构不影响推荐结果,但解释可信度随数据量下降

大型语言模型(LLMs)正被用于协同决策与生成解释,作为群体推荐系统(GRS)的联合决策者。本文通过对比社会选择理论中的聚合策略,评估其推荐与解释。结果显示,LLM推荐常与加性功利(ADD)聚合一致,但解释多提及评分平均化,虽相关却不完全对应。群体结构(统一或分歧)对推荐无影响。同时,模型频繁引入用户/物品相似性、多样性等未明确定义的标准,或使用未知流行度阈值。在更大项目集规模下,额外标准出现频率上升,暗示传统聚合方法效率不足。不一致且模糊的解释削弱了透明性与可解释性,而这两点正是使用LLM的核心动因。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly being implemented as joint decision-makers and explanation generators for Group Recommender Systems (GRS). In this paper, we evaluate these recommendations and explanations by comparing them to social choice-based aggregation strategies. Our results indicate that LLM-generated recommendations often resembled those produced by Additive Utilitarian (ADD) aggregation. However, the explanations typically referred to averaging ratings (resembling but not identical to ADD aggregation). Group structure, uniform or divergent, did not impact the recommendations. Furthermore, LLMs regularly claimed additional criteria such as user or item similarity, diversity, or used undefined popularity metrics or thresholds. Our findings have important implications for LLMs in the GRS pipeline as well as standard aggregation strategies. Additional criteria in explanations were dependent on the number of ratings in the group scenario, indicating potential inefficiency of standard aggregation methods at larger item set sizes. Additionally, inconsistent and ambiguous explanations undermine transparency and explainability, which are key motivations behind the use of LLMs for GRS.

群体推荐LLM解释可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。