构建首个大规模高质量生成模型与数据卡评估基准
MetaGAI: A Large-Scale and High-Quality Benchmark for Generative AI Model and Data Card Generation

- 采用多智能体框架,从论文、代码库、模型平台三源验证生成文档
- 包含2541组真实文档对,支持自动化与人工联合评估
- 适合研究自动文档生成、AI治理及模型可解释性的学者使用
生成式AI的快速发展迫切需要透明化和规范化文档标准。然而,手动制作模型与数据卡难以规模化,而现有自动化方法缺乏大规模、高保真度的评估基准。我们提出MetaGAI,一个涵盖2541个经验证的文档三元组的综合性基准,通过学术论文、GitHub仓库和Hugging Face资源的语义三角验证构建而成。不同于以往单一来源数据集,MetaGAI采用多智能体框架,包括检索、生成与编辑三个专用代理,并通过四维人机协同评估进行验证,包含人类对编辑后真实样本的评估。我们建立了一套结合自动化指标与经验证的LLM-as-a-Judge评估框架的稳健评测协议。大量分析表明,稀疏的混合专家架构在成本-质量效率上表现更优,且存在忠实度与完整性之间的根本权衡。MetaGAI为大规模自动化模型与数据卡生成方法的评测、训练与分析提供了基础测试平台。数据与代码已公开:https://github.com/haoxuan-unt2024/MetaGAI-Benchmark。
原文摘要 · Abstract (English)
The rapid proliferation of Generative AI necessitates rigorous documentation standards for transparency and governance. However, manual creation of Model and Data Cards is not scalable, while automated approaches lack large-scale, high-fidelity benchmarks for systematic evaluation. We introduce MetaGAI, a comprehensive benchmark comprising 2,541 verified document triplets constructed through semantic triangulation of academic papers, GitHub repositories, and Hugging Face artifacts. Unlike prior single-source datasets, MetaGAI employs a multi-agent framework with specialized Retriever, Generator, and Editor agents, validated through four-dimensional human-in-the-loop assessment, including human evaluation of editor-refined ground truth. We establish a robust evaluation protocol combining automated metrics with validated LLM-as-a-Judge frameworks. Extensive analysis reveals that sparse Mixture-of-Experts architectures achieve superior cost-quality efficiency, while a fundamental trade-off exists between faithfulness and completeness. MetaGAI provides a foundational testbed for benchmarking, training, and analyzing automated Model and Data Card generation methods at scale. Our data and code are available at: https://github.com/haoxuan-unt2024/MetaGAI-Benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。