arXiv:2510.26025cs.LG2025-10被引 3

用国际象棋测试AI是否真懂人类概念,发现越强的模型越不理解人类思维。

Exploring Human-AI Conceptual Alignment through the Prism of Chess

  • 用270M参数模型分析棋局,发现早期层能准确识别中心控制等概念
  • 深层网络性能虽高,但对人类概念的识别准确率降至50-65%
  • 随机开局数据集显示模型依赖记忆而非抽象理解,适合研究人机协作

AI系统是否真正理解人类概念,还是仅模仿表面模式?我们通过国际象棋这一兼具人类创造力与精确战略性的领域进行探究。分析一个参数量达2.7亿、达到特级大师水平的Transformer模型,发现显著矛盾:早期层对中心控制、马位等人类概念的编码准确率可达85%,而更深层网络虽提升整体表现,却逐渐偏离人类认知,准确率下降至50%-65%。为检验概念理解的鲁棒性,我们构建首个Chess960数据集——包含240个专家标注位置,覆盖6种战略概念。当通过随机起始布局消除开局理论影响后,所有方法的概念识别率均下降10%-20%,揭示模型依赖记忆而非抽象理解。层间分析表明,当前架构存在根本矛盾:赢得比赛的表示方式与人类思维不再对齐。这提示,随着性能优化,AI可能发展出越来越异化的智能,这对需要真正人机协作的创造性应用构成严峻挑战。数据集与代码见:https://github.com/slomasov/ChessConceptsLLM。

原文摘要 · Abstract (English)

Do AI systems truly understand human concepts or merely mimic surface patterns? We investigate this through chess, where human creativity meets precise strategic concepts. Analyzing a 270M-parameter transformer that achieves grandmaster-level play, we uncover a striking paradox: while early layers encode human concepts like center control and knight outposts with up to 85\% accuracy, deeper layers, despite driving superior performance, drift toward alien representations, dropping to 50-65\% accuracy. To test conceptual robustness beyond memorization, we introduce the first Chess960 dataset: 240 expert-annotated positions across 6 strategic concepts. When opening theory is eliminated through randomized starting positions, concept recognition drops 10-20\% across all methods, revealing the model's reliance on memorized patterns rather than abstract understanding. Our layer-wise analysis exposes a fundamental tension in current architectures: the representations that win games diverge from those that align with human thinking. These findings suggest that as AI systems optimize for performance, they develop increasingly alien intelligence, a critical challenge for creative AI applications requiring genuine human-AI collaboration. Dataset and code are available at: https://github.com/slomasov/ChessConceptsLLM.

人机协作概念理解国际象棋AI模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。