arXiv:2504.07801cs.IRcs.AI2025-04被引 10

用人格特质评估大模型推荐公平性,发现主流模型仍有显著偏差

FairEval: Evaluating Fairness in LLM-Based Recommendations with Personality Awareness

  • 将人格与8类敏感属性结合,系统评估推荐公平性
  • 测试中最大偏差达34.79%,两模型公平性得分超0.99
  • 适合关注算法公平性的推荐系统研究者和开发者

大语言模型在推荐系统中的应用日益广泛,但其在人口统计与心理特征维度的公平性仍存疑。我们提出FairEval,一个新型评估框架,用于系统分析基于大模型的推荐公平性。该框架融合人格特质与八类敏感人口属性(包括性别、种族、年龄等),实现用户层面的全面偏见评估。我们在音乐与电影推荐任务中评估了ChatGPT 4o与Gemini 1.5 Flash等模型。FairEval的公平性指标PAFS显示,ChatGPT 4o得分最高达0.9969,Gemini 1.5 Flash达0.9997,但最大偏差仍高达34.79%。结果凸显提示敏感性鲁棒性的重要性,为构建更具包容性的推荐系统提供支持。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) have enabled their application to recommender systems (RecLLMs), yet concerns remain regarding fairness across demographic and psychological user dimensions. We introduce FairEval, a novel evaluation framework to systematically assess fairness in LLM-based recommendations. FairEval integrates personality traits with eight sensitive demographic attributes,including gender, race, and age, enabling a comprehensive assessment of user-level bias. We evaluate models, including ChatGPT 4o and Gemini 1.5 Flash, on music and movie recommendations. FairEval's fairness metric, PAFS, achieves scores up to 0.9969 for ChatGPT 4o and 0.9997 for Gemini 1.5 Flash, with disparities reaching 34.79 percent. These results highlight the importance of robustness in prompt sensitivity and support more inclusive recommendation systems.

推荐系统公平性大模型人格分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。