研究大模型推荐中语言方言带来的偏见,发现不同方言影响推荐结果。
An Investigation of Linguistic Biases in LLM-Based Recommendations

- 用不同方言提示词测试大模型冷启动推荐表现
- 印度英语和混语提示导致推荐类型显著变化
- 适合关注AI公平性与多语言应用的研究者
本研究考察了在使用南美英语(AE)、印度英语(IE)及印英混用方言提示时,基于大语言模型的餐厅与商品推荐中的语言偏见。基于Yelp开放数据集(Yelp Inc., 2023)和Walmart商品评论数据集(PromptCloud, 2020),在冷启动设置下对大模型进行零样本提示,从按菜系与品类平衡的列表中选出前20项推荐。通过20组种子随机采样提升泛化性,聚合每类菜系与品类的响应计数,并对每个模型家族与主题(餐厅/商品)运行混合效应回归模型,采用似然比检验固定效应并进行事后成对比较。结果显示,方言显著影响推荐结果:Mistral-small-3.1及Llama-3.1系列模型对印度英语和混用提示更敏感;在商品推荐中,Llama-3.1-70B模型在七类中的四类对混用提示特别敏感;较大模型在印度英语提示下更多推荐美妆与家居类商品,较小模型则在混用提示下出现类似趋势。模型规模差异无统一规律,推荐结果随方言类型而异。
原文摘要 · Abstract (English)
We investigate linguistic biases in LLM-based restaurant and product recommendations given prompts varying across Southern American English (AE), Indian English (IE), and Code-Switched Hindi-English dialects, using the Yelp Open dataset (Yelp Inc., 2023) and Walmart product reviews dataset (PromptCloud,2020). We add lists of restaurant and product names balanced by cuisine type and product category to the prompts given to the LLM, and we zero-shot prompt the LLMs in a cold-start setting to select the top-20 restaurant and product recommendations from these lists for each of the dialect-varied prompts. We prompt LLMs using different list samples across 20 seeds for better generalization, and aggregate per cuisine-type and per category response counts for each seed, question/prompt, and LLM model. We run mixed-effects regression models for each model family and topic (restaurant/product) with the aggregate response counts as the dependent, and conduct likelihood ratio tests for the fixed effects with post-hoc pairwise testing of estimated marginal means differences, to investigate group-level differences in recommendation counts by model size and dialect type. Results show that dialect plays a role in the type of restaurant selected across the models tested with the mistral-small-3.1 model and both the llama-3.1 family models tested showing more sensitivity to Indian English and Code-Switched prompts. In terms of product recommendations, the llama-3.1-70B-model is particularly sensitive to Code-Switched prompts in four out of seven categories, and more beauty and home category recommendations are seen when using the Indian English and Code-Switched prompts for larger and smaller models, respectively. No broad trends are seen in the model-size based differences, with differing recommendations based on model sizes conditioned by the type of dialect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。