测试大模型能否零样本解决认知框架与符号接地难题。
Evaluating Large Language Models on the Frame and Symbol Grounding Problems: A Zero-shot Benchmark
- 设计两个零样本基准任务,模拟哲学中的框架问题与符号接地问题。
- 闭源模型在五次测试中表现稳定,得分显著高于开源模型。
- 适合关注AI认知能力边界与语言模型推理潜力的研究者。
近期大语言模型(LLMs)的发展重新点燃了人工智能领域的哲学讨论。框架问题与符号接地问题作为传统符号主义AI中长期无法解决的两大根本挑战,如今被重新审视。本研究通过设计两个反映各自哲学核心的基准任务,在零样本条件下对13个主流大模型(含闭源与开源)进行测试,每模型执行五次评估。响应结果从上下文推理、语义连贯性、信息过滤等多个维度评分。结果显示,尽管开源模型因参数规模、量化方式和指令微调差异表现出性能波动,但部分闭源模型始终获得高分。这表明某些现代大语言模型可能已具备应对这些长期理论难题的认知能力。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have revitalized philosophical debates surrounding artificial intelligence. Two of the most fundamental challenges - namely, the Frame Problem and the Symbol Grounding Problem - have historically been viewed as unsolvable within traditional symbolic AI systems. This study investigates whether modern LLMs possess the cognitive capacities required to address these problems. To do so, I designed two benchmark tasks reflecting the philosophical core of each problem, administered them under zero-shot conditions to 13 prominent LLMs (both closed and open-source), and assessed the quality of the models' outputs across five trials each. Responses were scored along multiple criteria, including contextual reasoning, semantic coherence, and information filtering. The results demonstrate that while open-source models showed variability in performance due to differences in model size, quantization, and instruction tuning, several closed models consistently achieved high scores. These findings suggest that select modern LLMs may be acquiring capacities sufficient to produce meaningful and stable responses to these long-standing theoretical challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。