arXiv:2505.13794cs.AI2025-05KDD被引 3

用大模型提取专家评估标准,让生态模型评价更智能可解释。

LLM-based Evaluation Policy Extraction for Ecological Modeling

  • 结合度量学习与大模型生成自然语言评估策略
  • 在多个数据集上准确捕捉碳通量预测的评估偏好
  • 适合需要可解释评价的生态建模研究者使用

生态时间序列评估对温室气体通量预测、碳氮循环模拟和水文周期监测等应用至关重要。传统数值指标(如R平方、均方根误差)虽广泛使用,但常无法捕捉生态过程特有的时间模式,需依赖专家视觉检查,耗时且难规模化。为此,我们提出一种新框架,融合度量学习与大语言模型(LLM)的自然语言策略提取,构建可解释的评估标准。该方法处理成对标注,通过策略优化机制生成并组合评估指标。在作物总初级生产力与二氧化碳通量预测的多个数据集上验证,结果表明其能有效捕捉合成与专家标注的评估偏好。该框架弥合了数值指标与专家知识间的鸿沟,提供可适应不同生态建模需求的可解释评估策略。

原文摘要 · Abstract (English)

Evaluating ecological time series is critical for benchmarking model performance in many important applications, including predicting greenhouse gas fluxes, capturing carbon-nitrogen dynamics, and monitoring hydrological cycles. Traditional numerical metrics (e.g., R-squared, root mean square error) have been widely used to quantify the similarity between modeled and observed ecosystem variables, but they often fail to capture domain-specific temporal patterns critical to ecological processes. As a result, these methods are often accompanied by expert visual inspection, which requires substantial human labor and limits the applicability to large-scale evaluation. To address these challenges, we propose a novel framework that integrates metric learning with large language model (LLM)-based natural language policy extraction to develop interpretable evaluation criteria. The proposed method processes pairwise annotations and implements a policy optimization mechanism to generate and combine different assessment metrics. The results obtained on multiple datasets for evaluating the predictions of crop gross primary production and carbon dioxide flux have confirmed the effectiveness of the proposed method in capturing target assessment preferences, including both synthetically generated and expert-annotated model comparisons. The proposed framework bridges the gap between numerical metrics and expert knowledge while providing interpretable evaluation policies that accommodate the diverse needs of different ecosystem modeling studies.

生态建模大模型评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。