对比人类与大模型对数学问题有趣性的判断差异。
A Matter of Interest: Understanding Interestingness of Math Problems in Humans and Language Models
- 用真实人类与模型评分对比,分析趣味性判断一致性。
- 模型与人类在趣味性分布上匹配度低,理由关联弱。
- 经筛选后,模型能生成有效且吸引人的数学题。
数学的发展很大程度上由问题的'有趣性'驱动:研究者和学生基于对趣味性和挑战性的预期选择研究或参与的问题。随着大型语言模型(LLMs)在数学研究与教育中日益普及,其判断是否与不同背景的人类一致变得至关重要。本文通过对比两类人群——具有大学数学背景的众包参与者和国际数学奥林匹克竞赛选手——的评分,评估了多个主流LLM在数学问题有趣性判断上的表现。尽管多数模型总体上与人类观点一致,但在判断分布上显著偏离;它们与人类选择理由的相关性也较低。此外,我们测试了模型生成有趣问题的能力,发现经过有效性过滤后,模型能够产出具有吸引力的问题。结论强调需构建多模型与人类协作系统,凸显了大模型在数学推理中的潜力与局限。
原文摘要 · Abstract (English)
The evolution of mathematics is shaped importantly by interestingness: researchers choose which problems to pursue, and students choose which problems to engage with, based on expectations of interest and challenge. As AI systems, particularly large language models (LLMs) that operate flexibly over natural language and formal mathematics, are increasingly used in mathematics research and education, it becomes crucial to characterize how closely their judgments align with people from different mathematical backgrounds. We study whether LLMs align with human interestingness judgments by comparing LLM ratings with those of two populations, crowdsourced participants with college math experience and International Math Olympiad competitors. Although many LLMs broadly agree with human notions of interestingness, they largely fail to match the distribution of human judgments. They also weakly align with why humans find problems interesting, with low correlation to human-selected rationales. Finally, we evaluate LLMs' ability to generate interesting problems and find that, after filtering for validity, LLMs are able to generate engaging problems. We conclude with takeaways, including the need for multi-LLM human-AI collaborative systems, that highlight both the promise and current limits of LLMs as partners in mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。