测试大模型在冷启动推荐中的偏见,发现性别文化刻板印象普遍存在。
Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting
- 构建可配置的冷启动推荐公平性评测框架,支持多领域多属性测试。
- Gemma 3 和 Llama 3.2 在音乐、电影、高校推荐中均显性别与文化偏见。
- 模型规模与公平性呈非线性关系,小模型未必更公平。
大型语言模型(LLMs)因其通用能力被广泛用于推荐任务。尽管在丰富上下文场景下表现良好,但在冷启动场景(仅含年龄、性别或语言等有限信息)中,其行为可能因预训练阶段编码的社会偏见而引发公平性问题。本文提出一个专门针对零上下文推荐公平性的评测基准。所设计的模块化流程支持可配置的推荐领域和敏感属性,可对任意开源大模型进行系统性、灵活的公平性审计。通过对当前领先模型(Gemma 3 和 Llama 3.2)的评估,发现其在音乐、电影及高校推荐等多个领域均存在一致的偏见,包括性别化与文化刻板印象。此外,研究揭示模型规模与公平性之间存在非线性关系,凸显需进行更细致的分析。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used for recommendation tasks due to their general-purpose capabilities. While LLMs perform well in rich-context settings, their behavior in cold-start scenarios, where only limited signals such as age, gender, or language are available, raises fairness concerns because they may rely on societal biases encoded during pretraining. We introduce a benchmark specifically designed to evaluate fairness in zero-context recommendation. Our modular pipeline supports configurable recommendation domains and sensitive attributes, enabling systematic and flexible audits of any open-source LLM. Through evaluations of state-of-the-art models (Gemma 3 and Llama 3.2), we uncover consistent biases across recommendation domains (music, movies, and colleges) including gendered and cultural stereotypes. We also reveal a non-linear relationship between model size and fairness, highlighting the need for nuanced analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。