提出客观评估推荐解释质量的新指标,区分准确性和用户契合度。
Towards a Signal Detection Based Measure for Assessing Information Quality of Explainable Recommender Systems
- 基于信号检测理论,将解释质量分解为保真度与契合度两个维度。
- 通过四类不同信息质量场景验证,指标能有效区分解释优劣。
- 适合研究可解释推荐系统评估的学者和工程师参考。
可解释推荐系统在提供推荐结果的同时,也给出推理依据,日益受到关注。当前评估多聚焦整体推荐性能,极少关注解释本身的质量。现有方法依赖用户主观调研,仅考察代表性解释因素对用户认知的影响,而非解释内容的真实性。本文旨在填补这一空白,提出一种客观度量解释信息质量(Veracity)的方法。将Veracity分解为两个维度:保真度(Fidelity)衡量解释是否包含关于推荐项的准确信息;契合度(Attunement)评估解释是否反映目标用户的偏好。基于信号检测理论,分别确定两个维度的决策结果,并合并计算敏感性,作为最终的Veracity值。为验证该指标有效性,设计了四类不同信息质量的测试场景,结果表明该指标能有效捕捉解释质量差异。
原文摘要 · Abstract (English)
There is growing interest in explainable recommender systems that provide recommendations along with explanations for the reasoning behind them. When evaluating recommender systems, most studies focus on overall recommendation performance. Only a few assess the quality of the explanations. Explanation quality is often evaluated through user studies that subjectively gather users' opinions on representative explanatory factors that shape end-users' perspective towards the results, not about the explanation contents itself. We aim to fill this gap by developing an objective metric to evaluate Veracity: the information quality of explanations. Specifically, we decompose Veracity into two dimensions: Fidelity and Attunement. Fidelity refers to whether the explanation includes accurate information about the recommended item. Attunement evaluates whether the explanation reflects the target user's preferences. By applying signal detection theory, we first determine decision outcomes for each dimension and then combine them to calculate a sensitivity, which serves as the final Veracity value. To assess the effectiveness of the proposed metric, we set up four cases with varying levels of information quality to validate whether our metric can accurately capture differences in quality. The results provided meaningful insights into the effectiveness of our proposed metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。