用智能代理自动生成10万+高质量图文配对图示。
Feynman: Knowledge-Infused Diagramming Agent for Scalable Visual Designs
- 通过分析领域知识生成可执行的声明式代码,逐步优化图示设计。
- 产出超10万对图文一致的图示数据,支持视觉推理模型评估。
- 适合需要大规模视觉语言数据的研究者与开发者使用。
视觉设计是先进多模态AI系统的重要应用。提升这些系统需大规模高质量的视觉-语言数据。尽管互联网上图像和文本数据丰富,但富含知识且对齐良好的图文对却稀少。本文提出基于智能体Feynman的可扩展图示生成流水线。Feynman首先枚举领域特定的知识组件(“想法”),并基于这些想法进行代码规划。根据计划,Feynman将想法转化为简洁的声明式程序,并通过反馈迭代优化图示。最终,这些程序由Penrose图示系统渲染。Penrose基于优化的渲染方式在保持视觉语义的同时引入随机性,从而生成具有一致性和多样性的图示。Feynman能以极低成本和时间生成带有真实描述的文字说明的图示。我们利用Feynman构建了一个包含超过10万对对齐良好的图示-标题数据集,并从新生成的数据中整理出一个视觉-语言基准测试集Diagramma,可用于评估视觉语言模型的视觉推理能力。我们计划将数据集、基准测试和完整智能体流水线开源发布。
原文摘要 · Abstract (English)
Visual design is an essential application of state-of-the-art multi-modal AI systems. Improving these systems requires high-quality vision-language data at scale. Despite the abundance of internet image and text data, knowledge-rich and well-aligned image-text pairs are rare. In this paper, we present a scalable diagram generation pipeline built with our agent, Feynman. To create diagrams, Feynman first enumerates domain-specific knowledge components (''ideas'') and performs code planning based on the ideas. Given the plan, Feynman translates ideas into simple declarative programs and iterates to receives feedback and visually refine diagrams. Finally, the declarative programs are rendered by the Penrose diagramming system. The optimization-based rendering of Penrose preserves the visual semantics while injecting fresh randomness into the layout, thereby producing diagrams with visual consistency and diversity. As a result, Feynman can author diagrams along with grounded captions with very little cost and time. Using Feynman, we synthesized a dataset with more than 100k well-aligned diagram-caption pairs. We also curate a visual-language benchmark, Diagramma, from freshly generated data. Diagramma can be used for evaluating the visual reasoning capabilities of vision-language models. We plan to release the dataset, benchmark, and the full agent pipeline as an open-source project.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。