揭示ChatGPT的美国化、单色化、顺性别偏见,直指训练数据是根源。
Stars, Stripes, and Silicon: Unravelling the ChatGPT's All-American, Monochrome, Cis-centric Bias
- 从训练数据质量与多样性出发,剖析大模型偏见成因
- 指出模型偏见会加剧社会不平等并生成有害内容
- 呼吁跨领域合作建立AI治理与问责机制
本文研究大型语言模型(如ChatGPT)中存在的偏见、毒性、不可靠性及鲁棒性不足问题。研究表明,这些问题主要源于训练数据的质量与多样性,而非模型架构本身。随着大模型在现实应用中日益普及,其放大既有偏见并生成有害内容的潜在危害已成为紧迫议题。论文呼吁开展跨学科协作,推动研究人员、从业者与利益相关方共同建立治理框架、监督机制与问责体系,以缓解偏见模型带来的负面影响。通过主动应对这些挑战,人工智能界可在不延续有害偏见或加剧现有不平等的前提下,充分发挥大模型对社会发展的巨大潜力。
原文摘要 · Abstract (English)
This paper investigates the challenges associated with bias, toxicity, unreliability, and lack of robustness in large language models (LLMs) such as ChatGPT. It emphasizes that these issues primarily stem from the quality and diversity of data on which LLMs are trained, rather than the model architectures themselves. As LLMs are increasingly integrated into various real-world applications, their potential to negatively impact society by amplifying existing biases and generating harmful content becomes a pressing concern. The paper calls for interdisciplinary efforts to address these challenges. Additionally, it highlights the need for collaboration between researchers, practitioners, and stakeholders to establish governance frameworks, oversight, and accountability mechanisms to mitigate the harmful consequences of biased LLMs. By proactively addressing these challenges, the AI community can harness the enormous potential of LLMs for the betterment of society without perpetuating harmful biases or exacerbating existing inequalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。