用多智能体框架评估地理数据的FAIR合规性,提升评估一致性与可审计性。
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets

- 构建13个专用LLM智能体,结合元数据提取与证据审查,逐项评分。
- 可复现性测试中评分一致率达89%,显著高于无审核机制的71%。
- 适合需高可信度数据评估的科研、政策制定者及数据治理团队。
地理空间数据支撑城市规划与气候建模等应用,但其FAIR合规性评估缺乏一致性。现有工具使用不同标准和证据源,对JavaScript渲染页面或特定仓库标识符表现不佳。在10个仓库的50个数据集上,各工具归一化得分的标准差平均为15.0个百分点,单个数据集最高达30.3。这些差异反映评估分歧而非准确度对比。本文提出AgentFAIR,一个融合结构化元数据提取与13个子原则专用LLM评估器的多智能体框架。每个评估器输出0-3分成熟度、引用证据与改进建议;批判者检查证据一致性并可触发针对性重评。平均可发现性、可访问性、互操作性、可重用性得分为79.7%、70.4%、45.3%、72.0%。与四种基线工具的相关系数为0.31至0.61,与FAIR-enough对比无统计显著性。在10个数据集的重复运行测试中,子原则一致性均值达89%(标准差3个百分点),未设批判者时仅为71%。初步15个数据集专家研究显示Fleiss' kappa为0.71,与专家共识一致率达82%。每数据集API成本约0.054美元。结果支持评估的可审计性与可行性,但受限于基准不全、缺失消融实验及单一模型族验证,对准确性与泛化能力的宣称仍有限。
原文摘要 · Abstract (English)
Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult. Existing evaluators use different rubrics and evidence sources and may fail on JavaScript-rendered pages or repository-specific identifiers. For 50 datasets from 10 repositories, the standard deviation of normalized scores across available tools averages 15.0 percentage points and reaches 30.3 for one dataset. Because these outputs are not equivalent measurements, we use them to characterize disagreement and failure modes, not comparative accuracy. We present AgentFAIR, a multi-agent framework combining structured metadata extraction with 13 sub-principle-specific LLM evaluators. Each produces a 0-3 maturity score, cited evidence, and recommendations; a critic checks evidence and consistency and can request targeted re-evaluation. Mean Findability, Accessibility, Interoperability, and Reusability scores are 79.7%, 70.4%, 45.3%, and 72.0%. Rank correlations with four baseline tools range from 0.31 to 0.61; the FAIR-enough comparison is not statistically significant. On a 10-dataset repeated-run subset, sub-principle agreement averages 89% (standard deviation: 3 percentage points), versus 71% without the critic. A preliminary 15-dataset expert study yields Fleiss' kappa of 0.71 and 82% alignment with expert consensus. API cost is approximately USD 0.054 per dataset. These results support auditability and feasibility, while the limited benchmark, incomplete ablations, and single-model-family validation constrain claims about accuracy and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。