打造跨学科科学智能评估工具,专测模型的科研能力。
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
- 聚焦科学核心能力,构建多模态与符号推理评测体系。
- 覆盖六大领域,基于真实科研数据设计任务挑战。
- 支持批量评估与自定义扩展,结果透明可复现。
我们推出SciEvalKit,一个统一的基准评估工具包,用于衡量AI模型在多个科学领域的通用智能水平。区别于通用评估平台,SciEvalKit专注于科学智能的核心能力,包括科学多模态感知、推理、理解、符号推理、代码生成、假说生成与科学知识理解。该工具包涵盖物理、化学、天文学、材料科学等六大主要科学领域,基于真实世界、领域特定的数据集构建专家级评测基准,确保任务反映真实的科学挑战。其灵活可扩展的评估流程支持模型与数据集的批量评估,兼容自定义集成,并提供透明、可复现、可比较的结果。通过融合能力导向评估与学科多样性,SciEvalKit为下一代科学基础模型与智能体提供标准化且可定制的评测基础设施。工具包已开源并持续维护,旨在推动社区驱动的AI for Science发展。
原文摘要 · Abstract (English)
We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the core competencies of scientific intelligence, including Scientific Multimodal Perception, Scientific Multimodal Reasoning, Scientific Multimodal Understanding, Scientific Symbolic Reasoning, Scientific Code Generation, Science Hypothesis Generation and Scientific Knowledge Understanding. It supports six major scientific domains, spanning from physics and chemistry to astronomy and materials science. SciEvalKit builds a foundation of expert-grade scientific benchmarks, curated from real-world, domain-specific datasets, ensuring that tasks reflect authentic scientific challenges. The toolkit features a flexible, extensible evaluation pipeline that enables batch evaluation across models and datasets, supports custom model and dataset integration, and provides transparent, reproducible, and comparable results. By bridging capability-based evaluation and disciplinary diversity, SciEvalKit offers a standardized yet customizable infrastructure to benchmark the next generation of scientific foundation models and intelligent agents. The toolkit is open-sourced and actively maintained to foster community-driven development and progress in AI4Science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。