arXiv:2604.03252cs.CYcs.CL2026-04

用大模型快速评估农业数字工具的包容性,替代人工耗时评审。

Evaluating Digital Inclusiveness of Digital Agri-Food Tools Using Large Language Models: A Comparative Analysis Between Human and AI-Based Evaluations

  • 用四款大模型对比人工评估农业工具的包容性
  • 部分维度下模型结果接近专家评分,但表现因模型而异
  • 适合资源有限或需快速评估的数字发展项目

确保数字包容性是农业食品系统中的关键任务,尤其在数字鸿沟依然存在的全球南方地区。多维数字包容性指数(MDII)提供了一个全面的人工主导框架,用于评估农业数字工具(agritools)的包容程度,但当前评估过程资源消耗大,通常需数月完成。本研究探讨大型语言模型(LLMs)是否能支持快速、基于AI的数字包容性评估,补充现有MDII工作流程。通过对比分析,研究基准了四种LLM(Grok、Gemini、GPT-4o和GPT-5)与以往专家评估的表现,考察模型评分与人类评分的一致性、温度设置敏感性及潜在偏见来源。结果表明,某些维度下,大模型生成的评估结果可近似专家判断,但可靠性在不同模型和情境间存在差异。这项探索性工作为生成式AI在包容性数字发展监测中的集成提供了早期证据,对时间敏感或资源受限环境下的评估规模化具有启示意义。

原文摘要 · Abstract (English)

Ensuring digital inclusiveness is a critical priority in agri-food systems, particularly in the Global South, where digital divides persist. The Multidimensional Digital Inclusiveness Index (MDII) offers a comprehensive, human-led framework to assess how inclusive digital agricultural tools (agritools) are. However, the current evaluation process is resource intensive, often requiring months to complete. This study explores whether large language models (LLMs) can support a rapid, AI-enabled assessment of digital inclusiveness, complementing the MDII's existing workflow. Using a comparative analysis, the research benchmarks the performance of four LLMs (Grok, Gemini, GPT-4o, and GPT-5) against prior expert-led evaluations. The study investigates model alignment with human scores, sensitivity to temperature settings, and potential sources of bias. Findings suggest that LLMs can generate evaluative outputs that approximate expert judgment in some dimensions, though reliability varies across models and contexts. This exploratory work provides early evidence for the integration of GenAI into inclusive digital development monitoring, with implications for scaling evaluations in time-sensitive or resource-constrained environments.

大模型评估数字包容农业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。