arXiv:2412.03605cs.CLcs.AI2024-12被引 15

揭示大模型认知偏差,用图谱找原因

CBEval: A framework for evaluating and interpreting cognitive biases in LLMs

  • 构建影响图谱,定位导致偏差的关键词句
  • 发现模型存在整数偏好和框架效应偏差
  • 适合研究模型可信性与公平性的学者

大语言模型(LLMs)在推理能力上迅速进步,但在认知过程仍存在显著缺陷。由于训练数据源自人类生成内容,这些模型可能继承认知偏差,影响其推理与决策可靠性。本文提出一个评估与解释认知偏差的框架,针对前沿语言模型展开研究,通过构建影响图谱,识别出导致偏差的关键短语与词汇。研究揭示了模型存在的整数偏好(round number bias)及框架效应(framing effect)所引发的认知壁垒,进一步阐明了偏差背后的生成机制。

原文摘要 · Abstract (English)

Rapid advancements in Large Language models (LLMs) has significantly enhanced their reasoning capabilities. Despite improved performance on benchmarks, LLMs exhibit notable gaps in their cognitive processes. Additionally, as reflections of human-generated data, these models have the potential to inherit cognitive biases, raising concerns about their reasoning and decision making capabilities. In this paper we present a framework to interpret, understand and provide insights into a host of cognitive biases in LLMs. Conducting our research on frontier language models we're able to elucidate reasoning limitations and biases, and provide reasoning behind these biases by constructing influence graphs that identify phrases and words most responsible for biases manifested in LLMs. We further investigate biases such as round number bias and cognitive bias barrier revealed when noting framing effect in language models.

大模型认知偏差可解释性推理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。