自动生成准确的AI基准文档,提升可比性和透明度。
Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
- 多智能体从异构源提取数据,大模型合成描述
- 用事实推理工具验证内容准确性,确保无误
- 适合需要快速评估基准的科研与工程人员
我们提出 Auto-BenchmarkCard,一种自动生成经验证的 AI 基准描述的工作流。基准文档常不完整或不一致,难以跨任务或领域进行解读与比较。该工作流结合多智能体从异构来源(如 Hugging Face、Unitxt、学术论文)中提取数据,并通过大语言模型进行合成。后续验证阶段使用 FactReasoner 工具,基于原子蕴含打分评估事实准确性。该流程有望提升 AI 基准报告的透明性、可比性和可复用性,帮助研究者与从业者更高效地选择和评估基准。
原文摘要 · Abstract (English)
We present Auto-BenchmarkCard, a workflow for generating validated descriptions of AI benchmarks. Benchmark documentation is often incomplete or inconsistent, making it difficult to interpret and compare benchmarks across tasks or domains. Auto-BenchmarkCard addresses this gap by combining multi-agent data extraction from heterogeneous sources (e.g., Hugging Face, Unitxt, academic papers) with LLM-driven synthesis. A validation phase evaluates factual accuracy through atomic entailment scoring using the FactReasoner tool. This workflow has the potential to promote transparency, comparability, and reusability in AI benchmark reporting, enabling researchers and practitioners to better navigate and evaluate benchmark choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。