首个评估大模型全面理性的基准,揭示其与人类理性差距
Rationality Check! Benchmarking the Rationality of Large Language Models
- 构建涵盖多领域的理性评测基准,覆盖理论与实践理性
- 发现大模型在复杂决策中表现显著偏离人类理性标准
- 适合模型开发者、AI伦理研究者及应用安全评估人员
大型语言模型(LLMs)作为深度学习与机器智能的最新进展,展现出惊人的能力,被认为是实现通用人工智能最有前景的方向之一。随着具备类人能力,LLMs被用于模拟人类并作为各类应用中的AI助手。因此,人们开始关注:在何种情况下LLMs会像真实人类一样思考和行动。理性是评估人类行为的核心概念,包括理论理性(思维层面)与实践理性(行动层面)。本文首次提出一个全面评估LLMs综合理性的基准,涵盖广泛领域与多种模型。该基准包含易用工具包、大量实验结果及分析,揭示了大模型在哪些方面趋近、哪些方面偏离理想人类理性。我们认为该基准可成为开发和使用大模型的基础工具。
原文摘要 · Abstract (English)
Large language models (LLMs), a recent advance in deep learning and machine intelligence, have manifested astonishing capacities, now considered among the most promising for artificial general intelligence. With human-like capabilities, LLMs have been used to simulate humans and serve as AI assistants across many applications. As a result, great concern has arisen about whether and under what circumstances LLMs think and behave like real human agents. Rationality is among the most important concepts in assessing human behavior, both in thinking (i.e., theoretical rationality) and in taking action (i.e., practical rationality). In this work, we propose the first benchmark for evaluating the omnibus rationality of LLMs, covering a wide range of domains and LLMs. The benchmark includes an easy-to-use toolkit, extensive experimental results, and analysis that illuminates where LLMs converge and diverge from idealized human rationality. We believe the benchmark can serve as a foundational tool for both developers and users of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。