研究提示词如何影响大模型推荐学者,揭示偏见来源。
Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

- 设计多维度提示词(身份、语言、地域等)测试模型推荐差异
- 日本提示生成高准确但单一的名单,南非提示准确性较低
- 提示词设计与模型选择同样关键,需系统审计
大型语言模型正越来越多地用于学者推荐,影响学术界对专家的认知。现有审计多局限于英语语境、单一学科且忽略提示词角色,难以理解输出差异来源。为此,我们构建基准,分离模型选择与提示设计对推荐结果的影响。通过改变提示词(语言、地理位置、角色与任务)和上下文(学科、资历、排名范围k),在六个科学领域中评估43个LLM的表现,将推荐结果与Semantic Scholar对比,衡量技术质量(事实性、覆盖度)与社会代表性(多样性、公平性)。结果显示:技术质量主要由模型决定,事实性与公平性受上下文影响,多样性则与地理位置相关。使用南非提示时推荐名单事实性较低,而日本提示产生高度准确但同质化、偏向高产学者的名单。因此,提示词设计是大模型学者发现中的非平凡因素,应与模型选择一同系统性审计。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as scholar recommenders, shaping who is seen as an expert in academia. Existing audits remain English-centric, single discipline, and persona-agnostic, leaving the source of output variability poorly understood. To this end, we propose a benchmark that disentangles the effects of model choice and prompt design on recommendations. We audit 43 LLMs by varying persona prompts (language, location, role-and-task) and context (field, seniority, k). Recommended scholars are compared against Semantic Scholar over six scientific disciplines to measure technical quality (factuality, coverage) and social representativeness (diversity, parity). Basic technical quality is driven by model choice, factuality and parity by context, and diversity by location. South Africa prompts yield less factual lists, while Japan prompts yield highly factual but homogeneous lists skewed toward highly productive scholars. Prompt design is thus a non-trivial axis of LLM-based scholar discovery and should be systematically audited alongside model choice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。