用程序生成几何图数据,让AI能准确理解几何题的图文描述。
Toward an Artificial General Teacher: Procedural Geometry Data Generation and Visual Grounding with Vision-Language Models
- 自动生成20万+张带精确分割掩码的几何图与多样描述。
- 微调后模型在几何图上达49% IoU,零样本仅<1%。
- 提出新评估指标缓冲IoU,更贴合细线结构定位需求。
我们将几何教育中的视觉解释问题视为指代图像分割(RIS):给定一张几何图和自然语言描述,需对所指几何元素生成像素级掩码。然而,现有基于自然图像基准(如RefCOCO)训练的RIS模型在几何图上表现极差,因摄影场景与抽象无纹理示意图存在根本域偏移。为解决训练数据缺失问题,我们提出全自动程序化数据生成引擎,可生成超过20万张合成几何图,包含像素级精确分割掩码与语言多样的指代表达,全程无需人工标注。我们进一步提出针对几何领域的视觉-语言模型(VLM)微调策略,结果表明微调后的Florence-2模型在标准IoU上达到49%,缓冲IoU(BIoU)达85%,而零样本设置下低于1%。我们引入缓冲IoU这一几何感知评估指标,考虑细结构定位误差,比标准IoU更能反映真实分割质量。本研究为构建能提供视觉化、分步解释的通用教师(AGT)奠定基础。
原文摘要 · Abstract (English)
We study visual explanation in geometry education as a Referring Image Segmentation (RIS) problem: given a diagram and a natural language description, the task is to produce a pixel-level mask for the referred geometric element. However, existing RIS models trained on natural image benchmarks such as RefCOCO fail catastrophically on geometric diagrams due to the fundamental domain shift between photographic scenes and abstract, textureless schematics. To address the absence of suitable training data, we present a fully automated procedural data engine that generates over 200,000 synthetic geometry diagrams with pixel-perfect segmentation masks and linguistically diverse referring expressions, requiring zero manual annotation. We further propose domain-specific fine-tuning of vision-language models (VLMs), demonstrating that a fine-tuned Florence-2 achieves 49% IoU and 85% Buffered IoU (BIoU), compared to <1% IoU in zero-shot settings. We introduce Buffered IoU, a geometry-aware evaluation metric that accounts for thin-structure localization, and show that it better reflects true segmentation quality than standard IoU. Our results establish a foundation for building Artificial General Teachers (AGTs) capable of providing visually grounded, step-by-step explanations of geometry problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。