分析12个大模型输出的相似性、多样性与偏见,揭示模型风格差异和伦理表现。
A Comprehensive Analysis of Large Language Model Outputs: Similarity, Diversity, and Bias
- 用5000个任务提示生成300万文本,对比12个主流模型输出特性。
- 同一模型输出更相似,GPT-4最多样,Llama3与Mistral最趋同。
- 发现部分模型性别平衡更好,为评估与改进模型伦理表现提供依据。
大型语言模型(LLMs)代表了迈向通用人工智能的重要进展,显著提升了人机交互能力。尽管它们在翻译、生成、编程和摘要等自然语言处理任务中表现优异,但其输出的相似性、多样性及伦理影响仍存疑问。例如,同一模型生成的文本有多相似?不同模型间有何差异?哪些模型更符合伦理标准?为此,我们使用5,000个涵盖生成、解释、重写等多样化任务的提示,生成了约300万条来自12个大模型的文本,包括OpenAI、Google、Microsoft、Meta和Mistral的开源与闭源系统。关键发现包括:(1)同一模型生成的文本彼此比人类写作更相似;(2)WizardLM-2-8x22b输出高度一致,而GPT-4生成内容更具变异性;(3)不同模型写作风格差异显著,Llama 3与Mistral相似度较高,GPT-4则表现出更强独特性;(4)词汇与语气差异凸显了大模型生成内容的语言独特性;(5)部分模型展现出更好的性别平衡与更低的偏见。这些结果为理解大模型输出行为与多样性提供了新视角,有助于未来模型开发与伦理评估。
原文摘要 · Abstract (English)
Large Language Models (LLMs) represent a major step toward artificial general intelligence, significantly advancing our ability to interact with technology. While LLMs perform well on Natural Language Processing tasks -- such as translation, generation, code writing, and summarization -- questions remain about their output similarity, variability, and ethical implications. For instance, how similar are texts generated by the same model? How does this compare across different models? And which models best uphold ethical standards? To investigate, we used 5{,}000 prompts spanning diverse tasks like generation, explanation, and rewriting. This resulted in approximately 3 million texts from 12 LLMs, including proprietary and open-source systems from OpenAI, Google, Microsoft, Meta, and Mistral. Key findings include: (1) outputs from the same LLM are more similar to each other than to human-written texts; (2) models like WizardLM-2-8x22b generate highly similar outputs, while GPT-4 produces more varied responses; (3) LLM writing styles differ significantly, with Llama 3 and Mistral showing higher similarity, and GPT-4 standing out for distinctiveness; (4) differences in vocabulary and tone underscore the linguistic uniqueness of LLM-generated content; (5) some LLMs demonstrate greater gender balance and reduced bias. These results offer new insights into the behavior and diversity of LLM outputs, helping guide future development and ethical evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。