arXiv:2512.24366cs.IR2025-12KDD被引 1

检测文本推荐解释的真假一致性,发现现有模型解释大多不靠谱。

On the Factual Consistency of Text-based Explainable Recommendation Models

  • 用大模型从评论中提取事实性陈述,构建可验证的基准数据
  • 实测六款顶尖模型,事实准确率仅4.38%~32.88%,远低于表面流畅度
  • 提出新评估方法,适合关注推荐系统可信度的研究者

基于文本的可解释推荐旨在生成自然语言解释以增强用户信任与系统透明度。尽管近期工作利用大模型生成流畅输出,但一个关键问题仍未被充分探讨:这些解释是否与已有证据在事实上一致?本文提出一个全面评估框架,设计基于提示的流水线,利用大模型从亚马逊评论中提取原子化解释语句,从而构建聚焦于事实内容的基准。该方法应用于亚马逊数据集五个类别,创建细粒度评估基准。进一步提出结合大模型与自然语言推理的语句级对齐指标,评估生成解释的事实一致性与相关性。在六种先进可解释推荐模型上进行广泛实验,发现虽语义相似度高(BERTScore F1: 0.81-0.90),但所有事实性指标表现极低(大模型语句级精确率:4.38%-32.88%)。结果凸显了在可解释推荐中引入事实感知评估的必要性,并为构建更可信的解释系统奠定基础。

原文摘要 · Abstract (English)

Text-based explainable recommendation aims to generate natural-language explanations that justify item recommendations, to improve user trust and system transparency. Although recent advances leverage LLMs to produce fluent outputs, a critical question remains underexplored: are these explanations factually consistent with the available evidence? We introduce a comprehensive framework for evaluating the factual consistency of text-based explainable recommenders. We design a prompting-based pipeline that uses LLMs to extract atomic explanatory statements from reviews, thereby constructing a ground truth that isolates and focuses on their factual content. Applying this pipeline to five categories from the Amazon Reviews dataset, we create augmented benchmarks for fine-grained evaluation of explanation quality. We further propose statement-level alignment metrics that combine LLM- and NLI-based approaches to assess both factual consistency and relevance of generated explanations. Across extensive experiments on six state-of-the-art explainable recommendation models, we uncover a critical gap: while models achieve high semantic similarity scores (BERTScore F1: 0.81-0.90), all our factuality metrics reveal alarmingly low performance (LLM-based statement-level precision: 4.38%-32.88%). These findings underscore the need for factuality-aware evaluation in explainable recommendation and provide a foundation for developing more trustworthy explanation systems.

可解释推荐事实一致性大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。