arXiv:2511.00076cs.LG2025-11被引 1

用贝塞尔曲线重建汉字,让AI学会理解字形的几何结构

Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves

  • 将汉字视为可执行的贝塞尔曲线程序,实现从图像到几何代码的逆向生成
  • 仅用现代汉字训练,就能零样本还原甲骨文,准确率超GPT-4o
  • 揭示模型掌握了抽象几何语法,超越了单纯像素识别

尽管视觉语言模型(VLMs)具备强大的语义理解能力,但其对视觉信息底层几何结构的解析能力仍待探索。象形文字兼具视觉形态与符号结构,是检验该能力的理想范例。本文将此视觉识别挑战置于数学领域,将每个字符表示为由几何基元构成的可执行程序。任务被建模为程序合成问题,训练VLM将位图图像反编译为由贝塞尔曲线组成的程序。我们的模型作为“视觉反编译器”,性能优于强零样本基线,包括GPT-4o。最显著的发现是:仅在现代汉字上训练的模型,可在零样本条件下成功重构古代甲骨文。这一泛化能力强烈表明,模型习得了抽象且可迁移的几何语法,实现了从像素级模式识别到更结构化视觉理解的跃迁。

原文摘要 · Abstract (English)

While Vision-language Models (VLMs) have demonstrated strong semantic capabilities, their ability to interpret the underlying geometric structure of visual information is less explored. Pictographic characters, which combine visual form with symbolic structure, provide an ideal test case for this capability. We formulate this visual recognition challenge in the mathematical domain, where each character is represented by an executable program of geometric primitives. This is framed as a program synthesis task, training a VLM to decompile raster images into programs composed of Bézier curves. Our model, acting as a "visual decompiler", demonstrates performance superior to strong zero-shot baselines, including GPT-4o. The most significant finding is that when trained solely on modern Chinese characters, the model is able to reconstruct ancient Oracle Bone Script in a zero-shot context. This generalization provides strong evidence that the model acquires an abstract and transferable geometric grammar, moving beyond pixel-level pattern recognition to a more structured form of visual understanding.

视觉理解几何建模字符重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。