评测大模型能否把自然语言几何题转成可执行的作图程序。
GeoBuildBench: A Benchmark for Interactive and Executable Geometry Construction from Natural Language

- 将几何题转为专用语言程序,需满足对象与约束条件。
- 489道中文教材题中模型常漏物或违反几何约束。
- 适合研究具身推理、多模态自修正能力的研究者。
我们提出GeoBuildBench,一个评估大语言模型和多模态智能体能否将非正式的自然语言平面几何问题转化为可执行几何构造的基准。与以往关注答案正确性或静态图解读的基准不同,GeoBuildBench将几何图视为交互式构造任务:给定文本问题,智能体需生成领域特定语言(DSL)程序,以生成满足明确指定几何对象和可验证约束的图。该基准包含489道经自动化筛选与人工验证的中文教材风格题目,确保文本完整且可构造。我们在受限迭代设置下评估多个前沿多模态模型,结果显示尽管成功率尚可,但模型频繁出现结构幻觉、遗漏对象及无法满足几何约束,且极少利用视觉和约束反馈进行自我修正。这些结果凸显几何构造作为超越文本或视觉合理性的具身可执行推理严谨测试床的价值。我们的基准与代码已公开。
原文摘要 · Abstract (English)
We introduce GeoBuildBench, a benchmark designed to evaluate whether large language models and multimodal agents can ground informal natural-language plane geometry problems into executable geometric constructions. Unlike existing geometry benchmarks that focus on answer correctness or static diagram interpretation, GeoBuildBench treats geometry diagram as an interactive construction task: given a textual problem, an agent must generate a domain-specific language (DSL) program to produce a diagram satisfying explicitly specified geometric objects and verifiable constraints. The benchmark features 489 Chinese textbook-style problems, curated through automated filtering and human validation to ensure text-complete, constructible problem specifications. We evaluate several state-of-the-art multimodal models in a bounded iterative setting and show that, despite reasonable success rates, models frequently exhibit structural hallucinations, missing objects, and failures to satisfy geometric constraints, with limited ability to exploit visual and constraint-based feedback for self-correction. These results highlight geometry construction as a rigorous testbed for grounded, executable reasoning beyond textual or visual plausibility. Our benchmark and code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。