为大模型推荐结果提供可信度评估与不确定性分解方法
Uncertainty Quantification and Decomposition for LLM-based Recommendation
- 提出量化推荐置信度的新框架,区分推荐与提示带来的不确定性
- 实验证明不确定性可有效反映推荐可靠性,且提示问题贡献主要不确定性
- 设计感知不确定性的提示策略,降低误差并提升推荐效果
尽管大型语言模型(LLMs)被广泛用于推荐系统,我们发现其推荐结果常存在不确定性。为确保LLM生成推荐的可信性,本文强调评估推荐可靠性的必要性。我们提出一种新框架,用于量化预测不确定性,以衡量基于LLM推荐的可靠性。进一步地,我们将预测不确定性分解为推荐不确定性与提示不确定性,从而深入分析不确定性的主要来源。通过大量实验,我们验证了:(1) 预测不确定性能有效指示推荐可靠性;(2) 利用分解后的不确定性度量揭示了不确定性的根源;(3) 提出一种不确定性感知的提示策略,可降低预测不确定性并增强推荐性能。代码与模型权重已公开于 https://github.com/WonbinKweon/UNC_LLM_REC_WWW2025。
原文摘要 · Abstract (English)
Despite the widespread adoption of large language models (LLMs) for recommendation, we demonstrate that LLMs often exhibit uncertainty in their recommendations. To ensure the trustworthy use of LLMs in generating recommendations, we emphasize the importance of assessing the reliability of recommendations generated by LLMs. We start by introducing a novel framework for estimating the predictive uncertainty to quantitatively measure the reliability of LLM-based recommendations. We further propose to decompose the predictive uncertainty into recommendation uncertainty and prompt uncertainty, enabling in-depth analyses of the primary source of uncertainty. Through extensive experiments, we (1) demonstrate predictive uncertainty effectively indicates the reliability of LLM-based recommendations, (2) investigate the origins of uncertainty with decomposed uncertainty measures, and (3) propose uncertainty-aware prompting for a lower predictive uncertainty and enhanced recommendation. Our source code and model weights are available at https://github.com/WonbinKweon/UNC_LLM_REC_WWW2025
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。