用视觉语言模型从图像生成植物结构模型,省去人工测量。
A Vision Language Model for Generating Procedural Plant Architecture Representations from Simulated Images
- 用合成图像训练视觉语言模型,从2D图生成3D植物架构的文本序列。
- 自回归生成时BLEU-4达94.00%,ROUGE-L为0.5182,可还原结构参数。
- 适合做植物数字化、计算机图形学或自动化植物建模的研究者。
三维程序化植物架构模型已成为模拟植物结构与功能、从实地测量中提取架构参数以及在计算机图形学中生成逼真植物的重要工具。然而,在田间尺度上测量这些模型的架构参数和嵌套结构仍极为耗时。本文提出一种新算法,仅通过图像生成3D植物架构,构建反映器官级几何与拓扑参数的功能性结构模型,提供更全面的植物架构表征。不同于使用3D传感器或多视角图像处理获取植物3D结构,本方法生成编码植物架构程序定义的标记序列。研究仅使用合成图像进行训练与测试,且已知精确的架构参数,以验证能否通过视觉语言模型(VLM)从图像数据中提取器官级架构参数。采用Helios 3D植物模拟器生成豇豆植物的合成数据集,植物架构信息以XML文件形式编码。开发了植物架构分词器,将XML文件转换为语言模型可预测的标记序列。教师强制训练下获得0.73的标记F1分数。自回归生成评估显示,BLEU-4得分为94.00%,ROUGE-L得分为0.5182。结果表明,从合成图像中生成植物架构模型并提取参数是可行的,未来工作将扩展至真实图像数据。
原文摘要 · Abstract (English)
Three-dimensional (3D) procedural plant architecture models have emerged as an important tool for simulation-based studies of plant structure and function, extracting plant architectural parameters from field measurements, and for generating realistic plants in computer graphics. However, measuring the architectural parameters and nested structures for these models at the field scales remains prohibitively labor-intensive. We present a novel algorithm that generates a 3D plant architecture from an image, creating a functional structural plant model that reflects organ-level geometric and topological parameters and provides a more comprehensive representation of the plant's architecture. Instead of using 3D sensors or processing multi-view images with computer vision to obtain the 3D structure of plants, we proposed a method that generates token sequences that encode a procedural definition of plant architecture. This work used only synthetic images for training and testing, with exact architectural parameters known, allowing testing of the hypothesis that organ-level architectural parameters could be extracted from image data using a vision-language model (VLM). A synthetic dataset of cowpea plant images was generated using the Helios 3D plant simulator, with the detailed plant architecture encoded in XML files. We developed a plant architecture tokenizer for the XML file defining plant architecture, converting it into a token sequence that a language model can predict. The model achieved a token F1 score of 0.73 during teacher-forced training. Evaluation of the model was performed through autoregressive generation, achieving a BLEU-4 score of 94.00% and a ROUGE-L score of 0.5182. This led to the conclusion that such plant architecture model generation and parameter extraction were possible from synthetic images; thus, future work will extend the approach to real imagery data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。