研究大模型推荐中的不确定性与公平性问题,揭示其偏见并提出新评估方法。
Uncertainty and Fairness Awareness in LLM-Based Recommendation Systems
- 构建包含8个人口属性的标注数据集,量化推荐不确定性与公平性差距。
- 发现Gemini 1.5 Flash在敏感属性上存在系统性不公平,相似度差距达0.1363。
- 引入性格特征感知的公平性评估,揭示个性化与群体公平的权衡关系。
大语言模型(LLMs)通过利用广泛上下文知识实现强大的零样本推荐,但预测不确定性与嵌入偏见威胁其可靠性与公平性。本文研究不确定性与公平性评估对LLM生成推荐的准确性、一致性及可信度的影响。我们引入一个经筛选的指标基准和一个标注了八个社会人口属性(31个类别值)的数据集,覆盖电影与音乐两个领域。通过深入案例研究,我们以熵为指标量化预测不确定性,并发现Google DeepMind的Gemini 1.5 Flash在某些敏感属性上表现出系统性不公平,相似度相关差距(SNSR)为0.1363,相似度方差(SNSV)为0.0507。这些偏差在提示扰动(如拼写错误、多语言输入)下依然存在。我们进一步将性格感知公平性融入RecLLM评估流程,揭示性格关联的偏见模式,并暴露个性化与群体公平之间的权衡。本文提出一种新型不确定性感知评估方法,提供深度不确定性案例的实证洞察,并引入基于性格画像的公平性基准,推动推荐系统的可解释性与公平性。这些贡献共同奠定更安全、可解释的RecLLM基础,激励未来多模型基准与自适应校准技术的发展。
原文摘要 · Abstract (English)
Large language models (LLMs) enable powerful zero-shot recommendations by leveraging broad contextual knowledge, yet predictive uncertainty and embedded biases threaten reliability and fairness. This paper studies how uncertainty and fairness evaluations affect the accuracy, consistency, and trustworthiness of LLM-generated recommendations. We introduce a benchmark of curated metrics and a dataset annotated for eight demographic attributes (31 categorical values) across two domains: movies and music. Through in-depth case studies, we quantify predictive uncertainty (via entropy) and demonstrate that Google DeepMind's Gemini 1.5 Flash exhibits systematic unfairness for certain sensitive attributes; measured similarity-based gaps are SNSR at 0.1363 and SNSV at 0.0507. These disparities persist under prompt perturbations such as typographical errors and multilingual inputs. We further integrate personality-aware fairness into the RecLLM evaluation pipeline to reveal personality-linked bias patterns and expose trade-offs between personalization and group fairness. We propose a novel uncertainty-aware evaluation methodology for RecLLMs, present empirical insights from deep uncertainty case studies, and introduce a personality profile-informed fairness benchmark that advances explainability and equity in LLM recommendations. Together, these contributions establish a foundation for safer, more interpretable RecLLMs and motivate future work on multi-model benchmarks and adaptive calibration for trustworthy deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。