构建农业多视角数据集,解决模型尺度混淆问题。
AgroOmni: A Large-Scale Multi-view Agricultural Dataset for Cross-Scale Multimodal Reasoning
- 构建28.8万张多尺度农业图像问答对,覆盖56类任务。
- 新模型在农业理解基准上达62.32%准确率,提升15.03%。
- 适合农业视觉与多模态研究者使用。
现代农业数据来自多种平台,涵盖从地面近景摄影到无人机航拍和卫星遥感影像的多尺度信息。因此,农业多模态推理需要强大的跨尺度空间理解能力。然而,由于缺乏多视角农业基准数据集,现有多模态大语言模型存在严重地面级偏差,导致尺度混淆和语义崩溃,例如将农田图像误识为墙壁或地板。为此,我们提出AgroOmni,一个包含28.8万组视觉问答对的大规模多视角训练语料库,覆盖14类任务中的56个专业任务类别,旨在捕捉现代精准农业中的多尺度多样性。基于该数据集,我们提出AgroNVILA,在AgroMind基准上达到62.32%的新最优成绩(相比GPT-5.2提升15.03%),有效缓解了多视图跨尺度差距,实现农业整体理解。在AgMMU上的诊断评估进一步揭示宏观先验与微观诊断间的固有异质性,且仅需微调即带来显著性能提升,充分证明了AgroOmni赋予模型的强大泛化能力。完整训练脚本已公开于https://anonymous.4open.science/r/AgroOmni-6510。
原文摘要 · Abstract (English)
Modern agricultural data is sourced from diverse platforms and spans multiple spatial scales, ranging from ground-level close-up photography to Unmanned Aerial Vehicle (UAV) aerial observation and satellite remote sensing imagery. Accordingly, agricultural multimodal reasoning demands robust cross-scale spatial understanding. However, due to the lack of multi-view agricultural benchmark datasets, existing multimodal large language models (MLLMs) exhibit severe ground-level bias, which leads to scale confusion then semantic collapse in agricultural perception tasks, such as misinterpreting farmland imagery as walls or floors. To address this, we introduce AgroOmni, a large-scale multi-view training corpus with 288K Visual Question Answering pairs covering 56 specialized task categories across 14 task types, designed to capture diverse scales in modern precision agriculture. Built on this dataset, we propose AgroNVILA, which achieves a new state-of-the-art of 62.32% on the AgroMind benchmark (+15.03% over GPT-5.2), effectively mitigating the multi-view cross-scale gap for holistic agricultural understanding. Diagnostic evaluations on AgMMU further reveal an inherent heterogeneity between macro-priors and micro-diagnostics through constrained zero-shot performance. Meanwhile, even minimal fine-tuning leads to a dramatic performance gain of AgroNVILA on AgMMU, strongly demonstrating its generalization capability empowered by AgroOmni. Full training scripts are publicly available at https://anonymous.4open.science/r/AgroOmni-6510.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。