首个融合语音视觉文本的多语言农业智能框架,解决数据与评估短板。
AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- 统一框架整合语音、视觉、文本模态,支持跨语言农业任务。
- 构建含492K合成+1.4K真实语音的六语种最大农业语音数据集。
- 推出首个三模态农业评测基准,支持多语言跨模态任务评估。
尽管多模态大模型发展迅速,农业应用仍受限于缺乏多语言语音数据、统一的多模态架构及全面的评估基准。为此,我们提出AgriGPT-Omni,一个集成语音、视觉与文本的农业通用框架。首先,构建可扩展的数据合成与采集管道,将农业文本与图像转化为训练数据,形成迄今最大的农业语音数据集,包含492K合成和1.4K真实语音样本,覆盖六种语言。其次,基于此数据,采用三阶段范式训练首个农业通用模型:文本知识注入、渐进式多模态对齐及基于GRPO的强化学习,实现跨语言与跨模态的统一推理。第三,提出AgriBench-Omni-2K,首个农业三模态基准,涵盖多样语音-视觉-文本任务及多语言子集,具备标准化协议与可复现工具。实验表明,AgriGPT-Omni在多语言与多模态推理及真实语音理解上显著优于通用基线。所有模型、数据、基准与代码将开源,以促进可复现研究、包容性农业智能及低资源地区可持续AI发展。
原文摘要 · Abstract (English)
Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the lack of multilingual speech data, unified multimodal architectures, and comprehensive evaluation benchmarks. To address these challenges, we present AgriGPT-Omni, an agricultural omni-framework that integrates speech, vision, and text in a unified framework. First, we construct a scalable data synthesis and collection pipeline that converts agricultural texts and images into training data, resulting in the largest agricultural speech dataset to date, including 492K synthetic and 1.4K real speech samples across six languages. Second, based on this, we train the first agricultural omni-model via a three-stage paradigm: textual knowledge injection, progressive multimodal alignment, and GRPO-based reinforcement learning, enabling unified reasoning across languages and modalities. Third, we propose AgriBench-Omni-2K, the first tri-modal benchmark for agriculture, covering diverse speech-vision-text tasks and multilingual slices, with standardized protocols and reproducible tools. Experiments show that AgriGPT-Omni significantly outperforms general-purpose baselines on multilingual and multimodal reasoning as well as real-world speech understanding. All models, data, benchmarks, and code will be released to promote reproducible research, inclusive agricultural intelligence, and sustainable AI development for low-resource regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。