让AI像编程一样逆向生成可编辑的图像,实现精准视觉重构。
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
- 通过符号逻辑与视觉感知交替验证,实现图像到可编辑代码的逆向生成。
- 在多个任务中准确率显著提升,最高达基准的124.70%。
- 无需训练即可支持文档、3D模型及物理交互等复杂任务,适合图形生成研究者。
视觉作为逆向图形(Vision-as-inverse-graphics)旨在将图像重构为可编辑的程序,但视觉语言模型(VLMs)在一次性设置下缺乏精细的空间定位能力,难以实现该目标。为此,我们提出VIGA(Vision-as-Inverse-Graphics Agent),一个交错式多模态推理框架,其中符号逻辑与视觉感知持续交叉验证。VIGA通过紧密耦合的代码-渲染-检查循环运行:合成符号程序,投影至视觉状态,并根据差异指导迭代修改。凭借高层语义能力和动态多模态记忆,VIGA可在长程任务中维持基于证据的持续优化。该无训练、任务无关框架无缝支持2D文档生成、3D重建、多步3D编辑及4D物理交互。最后,我们引入BlenderBench,一个具有挑战性的视觉到代码基准测试。实验表明,VIGA在BlenderGym(35.32%)、SlideBench(117.17%)和自建BlenderBench(124.70%)上显著优于单次基线方法。
原文摘要 · Abstract (English)
Vision-as-inverse-graphics, the concept of reconstructing images into editable programs, remains challenging for Vision-Language Models (VLMs), which inherently lack fine-grained spatial grounding in one-shot settings. To address this, we introduce VIGA (Vision-as-Inverse-Graphics Agent), an interleaved multimodal reasoning framework where symbolic logic and visual perception actively cross-verify each other. VIGA operates through a tightly coupled code-render-inspect loop: synthesizing symbolic programs, projecting them into visual states, and inspecting discrepancies to guide iterative edits. Equipped with high-level semantic skills and an evolving multimodal memory, VIGA sustains evidence-based modifications over long horizons. This training-free, task-agnostic framework seamlessly supports 2D document generation, 3D reconstruction, multi-step 3D editing, and 4D physical interaction. Finally, we introduce BlenderBench, a challenging visual-to-code benchmark. Empirically, VIGA substantially improves accuracy compared with one-shot baselines in BlenderGym (35.32%), SlideBench (117.17%) and our proposed BlenderBench (124.70%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。