arXiv:2410.16640cs.CL2024-10被引 1

用谚语测试大模型自评能力,发现其存在性别偏见和文化理解缺陷。

A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs

  • 构建300对语义相似但表述不同的谚语数据集,用于检测模型一致性。
  • 发现大模型在相似谚语间存在文本与数值不一致,暴露推理缺陷。
  • 适合关注AI伦理、偏见检测与模型可解释性的研究者参考。

大型语言模型(如ChatGPT、GPT-4、Claude-3、Llama)正广泛应用于各行业,但其推理与逻辑能力不足,引发对其作为评估工具的质疑。本文提出一种新的谚语推理任务,构建包含300对语义相近但表达不同谚语的数据集,涵盖性别、智慧与社会议题。通过测试文本一致性和数值一致性,验证了该方法能有效识别大模型在自评中的失败,揭示其在性别刻板印象和文化理解方面的缺陷。

原文摘要 · Abstract (English)

Large language models (LLMs) such as ChatGPT, GPT-4, Claude-3, and Llama are being integrated across a variety of industries. Despite this rapid proliferation, experts are calling for caution in the interpretation and adoption of LLMs, owing to numerous associated ethical concerns. Research has also uncovered shortcomings in LLMs' reasoning and logical abilities, raising questions on the potential of LLMs as evaluation tools. In this paper, we investigate LLMs' self-evaluation capabilities on a novel proverb reasoning task. We introduce a novel proverb database consisting of 300 proverb pairs that are similar in intent but different in wordings, across topics spanning gender, wisdom, and society. We propose tests to evaluate textual consistencies as well as numerical consistencies across similar proverbs, and demonstrate the effectiveness of our method and dataset in identifying failures in LLMs' self-evaluation which in turn can highlight issues related to gender stereotypes and lack of cultural understanding in LLMs.

大模型评估伦理偏见谚语推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。