构建数学协作讨论数据集,揭示大模型理解团队解题过程的短板
CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

- 收集164条专家标注的数学协作讨论链,追踪从问题到证明的演进过程
- 模型在预测下一条发言上准确率达83-88%,但角色识别仅0.42宏F1
- 适合研究协作推理、数学对话理解或人类-模型协同的学者
大型语言模型在数学推理方面取得显著进展,但现有基准多评估有明确答案的问题,缺乏对协作开放式问题求解的捕捉:参与者提出部分论证、发现先前步骤中的漏洞或错误、修复错误推理,并逐步整合增量贡献形成完整证明。我们引入CrowdMath,一个来自MIT PRIMES--AoPS CrowdMath项目(2016-2025)的164条专家标注进展链数据集,该协作研究计划已产出同行评审论文。每条链追踪多参与者的论坛讨论,从开放问题陈述推进至完整证明。帖子按其在演化解题过程中的功能角色标注,包括部分进展、证明完成、错误推理和错误识别。我们定义评估任务并测试六种前沿模型。模型在下一帖预测任务上达到83-88%准确率,表明能捕捉讨论的局部流。然而,在帖子角色分类上表现不佳,最佳模型仅得0.42宏F1。CrowdMath揭示了模型在解决规范问题与理解动态协作进展之间的差距。
原文摘要 · Abstract (English)
Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers, step-by-step solutions, or complete proofs. They do not capture collaborative open-problem solving: a setting in which participants propose partial arguments, identify gaps or errors in prior steps, repair flawed reasoning, and gradually synthesize incremental contributions into a proof. We introduce CrowdMath, a dataset of 164 expert-annotated progress chains from the MIT PRIMES--Art of Problem Solving (AoPS) CrowdMath program (2016-2025), a collaborative research initiative whose discussions have led to peer-reviewed publications. Each chain traces a multi-participant forum discussion from an open-problem statement to a completed proof. Posts are labeled by their functional roles in the evolving solution process, including partial progress, proof completion, erroneous reasoning, and error identification. We define evaluation tasks and benchmark six frontier models. Models achieve 83-88% accuracy on next-post prediction, suggesting that they can follow the local flow of mathematical discussion. However, they struggle to identify the functional significance of individual contributions with the best model achieving only 0.42 macro-F1 on post-role classification. CrowdMath exposes a gap between solving well-specified mathematical problems and understanding collaborative mathematical progress as it unfolds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。