arXiv:2509.04498cs.CLcs.AI2025-09中稿 · IJCNLP-AACL 2025 F…被引 4

评测大模型学术推荐中的偏见,发现北半球机构被过度推荐

Where Should I Study? Biased Language Models Decide! Evaluating Fairness in LMs for Academic Recommendations

  • 用360个虚拟用户测试3个开源模型的推荐倾向
  • 超2.5万条推荐显示全球南北差距与性别刻板印象
  • 提出多维评估框架,助力公平性改进

大型语言模型(LLMs)正被广泛用于教育规划等日常推荐系统,但其推荐可能延续社会偏见。本文实证研究了三个开源模型(LLaMA-3.1-8B、Gemma-7B、Mistral-7B)在大学和专业推荐中存在的人口统计、地理与经济偏见。基于360个涵盖性别、国籍与经济状况的模拟用户,分析超过25,000条推荐结果。结果显示:全球北方机构显著占优,推荐常强化性别刻板印象,且机构重复率高。尽管LLaMA-3.1-8B展现出最高多样性(推荐481所不同大学,覆盖58个国家),系统性不平等仍持续存在。为此,我们提出一种新型多维度评估框架,超越传统准确率,综合衡量人口与地理代表性。研究凸显在教育类大模型中重视偏见问题的紧迫性,以保障高等教育的全球公平获取。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used as daily recommendation systems for tasks like education planning, yet their recommendations risk perpetuating societal biases. This paper empirically examines geographic, demographic, and economic biases in university and program suggestions from three open-source LLMs: LLaMA-3.1-8B, Gemma-7B, and Mistral-7B. Using 360 simulated user profiles varying by gender, nationality, and economic status, we analyze over 25,000 recommendations. Results show strong biases: institutions in the Global North are disproportionately favored, recommendations often reinforce gender stereotypes, and institutional repetition is prevalent. While LLaMA-3.1 achieves the highest diversity, recommending 481 unique universities across 58 countries, systemic disparities persist. To quantify these issues, we propose a novel, multi-dimensional evaluation framework that goes beyond accuracy by measuring demographic and geographic representation. Our findings highlight the urgent need for bias consideration in educational LMs to ensure equitable global access to higher education.

大模型偏见评测教育推荐公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。