让两个模型从零开始通过沟通发明数学语言,检验其抽象推理能力。
Math Takes Two: A test for emergent mathematical reasoning in communication

- 两智能体在无先验知识下通过通信协作构建共享符号系统。
- 能自创数字体系的模型在视觉任务中表现更优,支持推理能力涌现。
- 适合研究语言与数学认知协同演化、具身智能的学者参考。
尽管语言模型在数学基准测试中表现优异,但其能力是否源于真正的数学推理,还是仅对形式语法的统计匹配仍不明确。现有评估多基于既定数学规范的符号问题,难以揭示模型从基本原理构建抽象概念的能力。本文提出「Math Takes Two」新基准,旨在通过通信检验数学推理的涌现。受人类数学认知与精准沟通共同演化的启发,该基准测试两个无数学背景的智能体能否通过协作发展出共享符号协议,以解决一个依赖数值系统实现外推的视觉任务。不同于多数数据集预设数学语言,本基准要求智能体从零发现潜在结构与表征。该设计为评估具备涌现数值推理能力的模型提供了新视角。
原文摘要 · Abstract (English)
Although language models demonstrate remarkable proficiency on mathematical benchmarks, it remains unclear whether this reflects true mathematical reasoning or statistical pattern matching over learning formal syntax. Most existing evaluations rely on symbolic problems grounded in established mathematical conventions, limiting insight into the models' ability to construct abstract concepts from first principles. In this work, we propose Math Takes Two, a new benchmark designed to assess the emergence of mathematical reasoning through communication. Motivated by the hypothesis that mathematical cognition in humans co-evolved with the need for precise communication, our benchmark tests whether two agents, without prior mathematical knowledge, can develop a shared symbolic protocol to solve a visually grounded task where the use of a numerical system facilitates extrapolation. Unlike many current datasets, our benchmark eschews predefined mathematical language, instead requiring agents to discover latent structure and representations from scratch. Math Takes Two thus provides a novel lens through which to develop and evaluate models with emergent numerical reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。