arXiv:2507.04224cs.CLcs.AI2025-07被引 3

评估大模型在学术图书馆服务中的公平性,发现其基本无歧视,仅一模型对女性有轻微刻板印象。

Fairness Evaluation of Large Language Models in Academic Library Reference Services

  • 用六款主流大模型测试不同性别、种族和身份用户的咨询响应差异。
  • 未发现种族/族裔差异,仅一个模型对女性有轻微刻板反应。
  • 模型根据用户身份调整正式度与专业术语,属合理适应而非歧视。

随着图书馆探索将大语言模型(LLMs)用于虚拟参考服务,一个核心问题浮现:这些模型能否平等服务于所有用户,不论其人口统计特征或社会地位?尽管它们具备可扩展支持的潜力,但可能复现训练数据中的社会偏见,威胁图书馆公平服务的承诺。为此,我们评估六款前沿大模型在面对性别、种族/族裔及机构角色不同的用户时,是否存在响应差异。结果显示,无证据表明存在种族或族裔差异;仅一款模型对女性表现出轻微刻板偏见。模型通过正式程度、礼貌性及领域专用词汇等语言选择,体现出对机构角色的细致回应,反映专业规范而非歧视性处理。研究结果表明,当前大模型在学术图书馆参考服务中已具备较高的公平性与情境适配能力。

原文摘要 · Abstract (English)

As libraries explore large language models (LLMs) for use in virtual reference services, a key question arises: Can LLMs serve all users equitably, regardless of demographics or social status? While they offer great potential for scalable support, LLMs may also reproduce societal biases embedded in their training data, risking the integrity of libraries' commitment to equitable service. To address this concern, we evaluate whether LLMs differentiate responses across user identities by prompting six state-of-the-art LLMs to assist patrons differing in sex, race/ethnicity, and institutional role. We find no evidence of differentiation by race or ethnicity, and only minor evidence of stereotypical bias against women in one model. LLMs demonstrate nuanced accommodation of institutional roles through the use of linguistic choices related to formality, politeness, and domain-specific vocabularies, reflecting professional norms rather than discriminatory treatment. These findings suggest that current LLMs show a promising degree of readiness to support equitable and contextually appropriate communication in academic library reference services.

大模型公平性图书馆偏见评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。