用大模型投票过滤数学知识库的噪声,提升概念分类准确性。
Categorizing Mathematical Concepts with LLM Voting Ensembles in Mathswitch

- 构建大模型投票集成系统,自动判断数学概念条目是否有效
- 在已知正确标签的数据上验证,准确率超85%且对上下文敏感
- 发现三类典型错误,为知识库纠错提供具体改进方向
Mathswitch 是一个开源项目,从 Wikidata、Wikipedia、MathWorld 等来源导入数学概念条目,并链接指向同一概念的不同记录。项目不重构或重新定义内容,各源保持原有结构。当前重点是通过 Wikidata 查询导入数据,并计划扩展至更多资源和增强概念关联。由于依赖协作编辑的图谱查询,导入数据存在噪声:部分条目非数学内容,部分描述模糊。本文测试大模型投票集成能否过滤此类噪声。在具有已知 MathWorld 标识符的 Wikidata 条目上进行评估作为正样本控制,研究当数据库标识符被移除时分类结果的变化。分析模型分歧案例,将其分为三类:描述过于简略、范围过窄偏差、编纂范围不匹配,提示不同修复策略。
原文摘要 · Abstract (English)
Mathswitch is an open-source project that imports mathematical concept records from sources such as Wikidata, Wikipedia, MathWorld, Encyclopedia of Mathematics, nLab, ProofWiki, and Agda-Unimath, and links records that refer to the same concept. It does not reorganize or redefine the imported content; each source retains its own structure. The current focus is on importing concept data from Wikidata and the resources it links to, with plans to expand to further sources and better concept linking. Because the concept set is approximated through queries over Wikidata's collaboratively edited graph, the imported data is noisy: some items are non-mathematical, while others are ambiguous. In this paper, we test whether a voting ensemble of LLM judges can filter this noise. We evaluate it on Wikidata items with known MathWorld identifiers as a positive control, and examine how classification changes when database identifiers are removed from context. We then inspect the cases where the judges disagree with MathWorld and group these disagreements into three categories (degenerate descriptions, narrow scope bias, and editorial-scope mismatches) that suggest different remediation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。