arXiv:2410.00175cs.CL2024-10EMNLP被引 4

大模型既能批判也能为性别歧视言论辩护,揭示其道德立场的多样性。

Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse

  • 分析8个大模型对性别歧视内容的道德解释,发现其立场可正可反。
  • 模型输出内容清晰且贴合语境,能反映不同意识形态对性别的看法。
  • 适合关注社会议题、伦理风险与模型安全的研究者参考。

本研究揭示了大语言模型在性别歧视内容上表现出的可变道德立场:既能批评,也能为性别偏见辩护。我们评估了八个大模型,发现它们均能基于不同的道德框架,生成合理且上下文相关的解释,涵盖从进步到保守的不同性别观念。通过人工与自动评估,所有模型输出均具可读性和情境相关性,有助于理解社会对性别歧视的认知差异。进一步分析显示,模型引用的道德基础反映了多元意识形态,部分模型更倾向特定性别角色立场。研究警示需警惕模型被滥用以正当化性别歧视,同时指出其在揭示性别偏见根源和设计干预措施方面的潜力。因此,在涉及敏感社会议题的应用中,必须加强对模型的监控与安全机制设计。

原文摘要 · Abstract (English)

This work provides an explanatory view of how LLMs can apply moral reasoning to both criticize and defend sexist language. We assessed eight large language models, all of which demonstrated the capability to provide explanations grounded in varying moral perspectives for both critiquing and endorsing views that reflect sexist assumptions. With both human and automatic evaluation, we show that all eight models produce comprehensible and contextually relevant text, which is helpful in understanding diverse views on how sexism is perceived. Also, through analysis of moral foundations cited by LLMs in their arguments, we uncover the diverse ideological perspectives in models' outputs, with some models aligning more with progressive or conservative views on gender roles and sexism. Based on our observations, we caution against the potential misuse of LLMs to justify sexist language. We also highlight that LLMs can serve as tools for understanding the roots of sexist beliefs and designing well-informed interventions. Given this dual capacity, it is crucial to monitor LLMs and design safety mechanisms for their use in applications that involve sensitive societal topics, such as sexism.

大模型伦理性别歧视道德推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。