DEEP自动化评估机器翻译与文字识别模型,支持Docker部署并可视化结果差异。
DEEP: Docker-based Execution and Evaluation Platform
- 基于Docker的自动化执行与评分框架,支持多模型并行测试
- 通过统计聚类识别模型性能分组,揭示显著性差异
- 提供可视化网页工具,帮助理解模型表现与结果意义
在研究中,系统间的对比评估是常见任务,对选择合适系统或展示研究成果至关重要,也是公开竞赛的核心环节。本文提出的DEEP软件可自动化执行并评分机器翻译与光学字符识别模型。该平台支持接收容器化系统(Dockerized),运行后自动提取信息并进行评估。通过基于评估指标的统计分析,采用聚类算法识别各模型性能的显著差异,帮助评估者理解不同模型之间的表现分组。此外,系统配套提供可视化网页应用,使结果更易解读。最后,文中展示了一个典型使用案例。
原文摘要 · Abstract (English)
Comparative evaluation of several systems is a recurrent task in researching. It is a key step before deciding which system to use for our work, or, once our research has been conducted, to demonstrate the potential of the resulting model. Furthermore, it is the main task of competitive, public challenges evaluation. Our proposed software (DEEP) automates both the execution and scoring of machine translation and optical character recognition models. Furthermore, it is easily extensible to other tasks. DEEP is prepared to receive dockerized systems, run them (extracting information at that same time), and assess hypothesis against some references. With this approach, evaluators can achieve a better understanding of the performance of each model. Moreover, the software uses a clustering algorithm based on a statistical analysis of the significance of the results yielded by each model, according to the evaluation metrics. As a result, evaluators are able to identify clusters of performance among the swarm of proposals and have a better understanding of the significance of their differences. Additionally, we offer a visualization web-app to ensure that the results can be adequately understood and interpreted. Finally, we present an exemplary case of use of DEEP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。