arXiv:2603.08930cs.CVcs.AI2026-03

用视觉语言模型从无人机图像生成植物模拟参数,提升农业数字孪生效率。

Using Vision Language Foundation Models to Generate Plant Simulation Configurations via In-Context Learning

  • 利用VLM直接从遥感图像生成JSON格式的植物模拟配置。
  • 模型能准确估计植株数量和太阳方位角,但受上下文偏差影响性能下降。
  • 适合农业数字孪生、智能种植系统研究者参考。

本文提出一个合成基准,用于评估视觉语言模型(VLMs)在生成农业数字孪生植物模拟配置方面的能力。尽管功能-结构植物模型(FSPMs)可用于模拟农业环境中的生物物理过程,但其高复杂性和低吞吐量限制了大规模部署。我们提出一种新方法,利用开源先进VLM(Gemma 3 和 Qwen3-VL),从基于无人机的遥感图像中直接生成JSON格式的模拟参数。使用通过Helios 3D程序化植物生成库创建的合成豇豆田数据集,测试了五种上下文学习方法,并在三类指标下评估模型表现:JSON完整性、几何评估与生物物理评估。结果表明,尽管VLM能理解结构元数据并估计植株数量与太阳方位角,但在视觉线索不足时易受上下文偏见影响或依赖数据集均值。在真实世界无人机正射影像数据集上的验证及盲基线消融实验进一步揭示了模型推理能力与其对上下文先验的依赖关系。据我们所知,这是首个利用VLM生成植物模拟结构化JSON配置的研究,为农业数字孪生中3D田块重建提供了可扩展框架。

原文摘要 · Abstract (English)

This paper introduces a synthetic benchmark to evaluate the performance of vision language models (VLMs) in generating plant simulation configurations for digital twins. While functional-structural plant models (FSPMs) are useful tools for simulating biophysical processes in agricultural environments, their high complexity and low throughput create bottlenecks for deployment at scale. We propose a novel approach that leverages state-of-the-art open-source VLMs -- Gemma 3 and Qwen3-VL -- to directly generate simulation parameters in JSON format from drone-based remote sensing images. Using a synthetic cowpea plot dataset generated via the Helios 3D procedural plant generation library, we tested five in-context learning methods and evaluated the models across three categories: JSON integrity, geometric evaluations, and biophysical evaluations. Our results show that while VLMs can interpret structural metadata and estimate parameters like plant count and sun azimuth, they often exhibit performance degradation due to contextual bias or rely on dataset means when visual cues are insufficient. Validation on a real-world drone orthophoto dataset and an ablation study using a blind baseline further characterize the models' reasoning capabilities versus their reliance on contextual priors. To the best of our knowledge, this is the first study to utilize VLMs to generate structural JSON configurations for plant simulations, providing a scalable framework for reconstruction 3D plots for digital twin in agriculture.

数字孪生植物建模视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。