评测6个大模型在物理学者推荐中的偏见与准确性,发现普遍存在性别、族裔和资历偏差。
Whose Name Comes Up? Auditing LLM-Based Scholar Recommendations
- 用真实学术数据对比大模型推荐结果,评估其一致性与事实性。
- 所有模型均存在偏见,70亿参数模型波动最大,多数推荐重复且格式错误。
- 严重偏向资深学者,男性、白人科学家被过度推荐,亚洲学者被低估。
本文评估六种开源大模型(llama3-8b、llama3.1-8b、gemma2-9b、mixtral-8x7b、llama3-70b、llama3.1-70b)在五个任务中的物理学者推荐表现:按领域推荐前k名专家、按学科/时代/资历推荐有影响力学者、以及学者对应者。基于美国物理学会和OpenAlex的真实数据建立基准,分析模型输出的一致性、事实性及性别、族裔、学术声望、学者相似性等偏见。结果表明所有模型均存在不一致与偏见:mixtral-8x7b输出最稳定,llama3.1-70b波动最大;多数模型出现重复推荐,gemma2-9b和llama3.1-8b存在严重格式错误。尽管推荐多为真实学者,但在领域、时代、资历特定查询中准确率下降,普遍青睐资深学者。性别偏见显著,反映男性主导;亚洲学者代表性不足,白人学者被高估。虽部分体现机构与合作网络多样性,但模型仍倾向高被引、高产出学者,强化‘马太效应’,地理分布有限。研究凸显改进模型以实现更可靠、公平学术推荐的必要性。
原文摘要 · Abstract (English)
This paper evaluates the performance of six open-weight LLMs (llama3-8b, llama3.1-8b, gemma2-9b, mixtral-8x7b, llama3-70b, llama3.1-70b) in recommending experts in physics across five tasks: top-k experts by field, influential scientists by discipline, epoch, seniority, and scholar counterparts. The evaluation examines consistency, factuality, and biases related to gender, ethnicity, academic popularity, and scholar similarity. Using ground-truth data from the American Physical Society and OpenAlex, we establish scholarly benchmarks by comparing model outputs to real-world academic records. Our analysis reveals inconsistencies and biases across all models. mixtral-8x7b produces the most stable outputs, while llama3.1-70b shows the highest variability. Many models exhibit duplication, and some, particularly gemma2-9b and llama3.1-8b, struggle with formatting errors. LLMs generally recommend real scientists, but accuracy drops in field-, epoch-, and seniority-specific queries, consistently favoring senior scholars. Representation biases persist, replicating gender imbalances (reflecting male predominance), under-representing Asian scientists, and over-representing White scholars. Despite some diversity in institutional and collaboration networks, models favor highly cited and productive scholars, reinforcing the rich-getricher effect while offering limited geographical representation. These findings highlight the need to improve LLMs for more reliable and equitable scholarly recommendations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。