用多智能体进化框架逆向生成几何图像代码,精准还原图形逻辑
Geo-Code: A Code Framework for Reverse Code Generation from Geometric Images Based on Two-Stage Multi-Agent Evolution
- 分两阶段:先像素锚定坐标,再通过视觉反馈迭代优化代码
- 重建准确率显著提升,重构图像在多模态任务中表现接近原图
- 开源超1500样本数据集和模型,降低研究门槛
程序代码是连接视觉与逻辑的桥梁,通过辅助线构造、视角变换等几何操作,可有效提升大模型的多模态推理能力。然而,现有逆图形方法在重建复杂几何细节时面临巨大挑战,常导致关键几何约束丢失或结构失真。为此,我们提出Geo-coder——首个基于多智能体系统的几何图像逆向编程框架。方法创新性地将过程解耦为:第一阶段利用视觉算子与大模型优势,精确捕捉像素坐标与视觉属性;第二阶段引入合成-渲染-验证闭环,双向视觉反馈驱动代码自我修正。大量实验表明,Geo-coder在几何重建精度与视觉一致性上均有显著提升。尤其值得注意的是,本方法有效保留核心几何语义,重构图像在多模态推理任务中的表现与原图相当,充分验证了框架的鲁棒性。最后,为降低研究成本,我们基于Geo-coder框架开源了包含超过1500个样本的Geo-coder数据集,并发布了GeocodeLM模型,为该领域后续研究奠定了坚实的数据与模型基础。
原文摘要 · Abstract (English)
Program code serves as a bridge linking vision and logic, providing a feasible supervisory approach for enhancing the multimodal reasoning capability of large models through geometric operations such as auxiliary line construction and perspective transformation. Nevertheless, current inverse graphics methods face tremendous challenges in accurately reconstructing complex geometric details, which often results in the loss of key geometric constraints or structural distortion. To address this bottleneck, we propose Geo-coder -- the first inverse programming framework for geometric images based on a multi-agent system. Our method innovatively decouples the process into geometric modeling via pixel-wise anchoring and metric-driven code evolution: Stage 1 leverages the complementary advantages of visual operators and large models to achieve precise capture of pixel coordinates and visual attributes; Stage 2 introduces a synthesis-rendering-validation closed loop, where bidirectional visual feedback drives the self-correction of code. Extensive experiments demonstrate that Geo-coder achieves a substantial lead in both geometric reconstruction accuracy and visual consistency. Notably, by effectively preserving the core geometric semantics, the images reconstructed with our method exhibit equivalent performance to the original ones in multimodal reasoning tasks, which fully validates the robustness of the framework. Finally, to further reduce research costs, we have open-sourced the Geo-coder dataset constructed on the GeoCode framework, which contains more than 1,500 samples. On this basis, we have also open-sourced the GeocodeLM model, laying a solid data and model foundation for subsequent research in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。