分析大模型生成的女性故事,发现黑人女性叙事隐含殖民偏见
Clustering Discourses: Racial Biases in Short Stories about Women Generated by Large Language Models
- 用聚类算法分析2100篇葡萄牙语短故事
- 发现三类典型叙事:社会逆袭、祖先神话与自我实现
- 结合机器学习与人工分析,揭示语言中的结构性偏见
本研究探讨大型语言模型(特别是LLaMA 3.2-3B)在葡萄牙语环境下生成关于黑人与白人女性的短故事时如何建构叙事。从2100篇生成文本中,采用计算方法对语义相似的故事进行聚类,进而选取样本进行定性分析。结果揭示出三种主要话语模式:社会克服、祖先神话化与主观自我实现。分析表明,尽管文本语法正确、表面中立,却固化了殖民结构下的女性身体观念,强化了历史不平等。研究提出一种融合机器学习与人工话语分析的综合方法,以系统识别和理解模型中的隐性偏见。
原文摘要 · Abstract (English)
This study investigates how large language models, in particular LLaMA 3.2-3B, construct narratives about Black and white women in short stories generated in Portuguese. From 2100 texts, we applied computational methods to group semantically similar stories, allowing a selection for qualitative analysis. Three main discursive representations emerge: social overcoming, ancestral mythification and subjective self-realization. The analysis uncovers how grammatically coherent, seemingly neutral texts materialize a crystallized, colonially structured framing of the female body, reinforcing historical inequalities. The study proposes an integrated approach, that combines machine learning techniques with qualitative, manual discourse analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。