发现大模型旅行推荐存在族裔与性别偏见,少数群体更易被误荐。
Whose Journey Matters? Investigating Identity Biases in Large Language Models (LLMs) for Travel Planning Assistance
- 用公平性探测分析三款主流开源大模型的旅行推荐
- 族裔和性别分类器准确率高于随机水平,存在刻板印象偏差
- 少数群体推荐中幻觉现象更频繁,影响推荐可靠性
随着大型语言模型(LLMs)在旅游与酒店行业的广泛应用,其对不同身份群体的服务公平性问题持续引发关注。基于社会认同理论与社会技术系统理论,本研究探究了大模型在旅行推荐中体现的族裔与性别偏见。通过公平性探测,分析了三款领先开源大模型的输出结果。结果显示,族裔与性别分类器的测试准确率均显著高于随机水平。对关键特征的分析揭示了大模型生成推荐中存在刻板印象偏差。此外,发现这些特征中存在幻觉现象,且在少数群体的推荐中更为频繁。研究表明,当作为旅行规划助手时,大模型表现出族裔与性别偏见。该研究强调需采取偏见缓解策略,以提升生成式AI驱动旅行推荐的包容性与可靠性。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly integral to the hospitality and tourism industry, concerns about their fairness in serving diverse identity groups persist. Grounded in social identity theory and sociotechnical systems theory, this study examines ethnic and gender biases in travel recommendations generated by LLMs. Using fairness probing, we analyze outputs from three leading open-source LLMs. The results show that test accuracy for both ethnicity and gender classifiers exceed random chance. Analysis of the most influential features reveals the presence of stereotype bias in LLM-generated recommendations. We also found hallucinations among these features, occurring more frequently in recommendations for minority groups. These findings indicate that LLMs exhibit ethnic and gender bias when functioning as travel planning assistants. This study underscores the need for bias mitigation strategies to improve the inclusivity and reliability of generative AI-driven travel planning assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。