arXiv:2507.01853cs.CL2025-07

一站式评估低资源多语言大模型,支持55+语种基准测试

Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models

  • 构建模块化框架,零代码网页+交互式命令行双接口
  • 覆盖9类任务,支持本地与私有模型,11项核心功能可插拔
  • 相比5个基线工具,用户满意度和设置效率提升2倍以上

大型语言模型的快速发展凸显了具备全球适用性、灵活性和模块化的评估框架的必要性。我们提出EKA-EVAL,一个统一的端到端框架,结合零代码网页界面和交互式命令行,确保广泛可访问性。该框架整合了涵盖九个评估类别的55个以上多语言基准测试,支持本地和专有模型,并通过模块化、即插即用的架构提供11项核心能力。设计用于可扩展的多语言评估,尤其支持低资源多语言场景。据我们所知,EKA-EVAL是首个在单一平台上实现全面覆盖的评估套件。与五个现有基线的对比显示,在关键可用性指标上至少提升2倍,用户满意度最高,设置时间更短,且基准测试结果一致可复现。该框架开源,可在https://github.com/lingo-iitgn/eka-eval获取。

原文摘要 · Abstract (English)

The rapid evolution of Large Language Models' has underscored the need for evaluation frameworks that are globally applicable, flexible, and modular, and that support a wide range of tasks, model types, and linguistic settings. We introduce EKA-EVAL, a unified, end- to-end framework that combines a zero-code web interface and an interactive CLI to ensure broad accessibility. It integrates 55+ multilingual benchmarks across nine evaluation categories, supports local and proprietary models, and provides 11 core capabilities through a modular, plug-and-play architecture. Designed for scalable, multilingual evaluation with support for low-resource multilingual languages, EKA-EVAL is, to the best of our knowledge, the first suite to offer comprehensive coverage in a single platform. Comparisons against five existing baselines indicate improvements of at least 2x better on key usability measures, with the highest user satisfaction, faster setup times, and consistent benchmark reproducibility. The framework is open-source and publicly available at https://github.com/lingo-iitgn/eka-eval.

多语言模型评估框架低资源语言开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。