2025年语音语言评估挑战赛,聚焦英语学习者口语评分与纠错。
Speak & Improve Challenge 2025: Tasks and Baseline Systems
- 构建多任务挑战赛,涵盖语音识别到语法纠错反馈
- 发布315小时二语英语口语数据集,含人工转录与错误标注
- 提供封闭/开放双赛道,适合语音与教育技术研究者
本文介绍与ISCA SLaTE 2025研讨会相关的「Speak & Improve Challenge 2025:口语语言评估与反馈」挑战赛。目标是推动口语语言评估与学习反馈的研究,涵盖底层技术与语言学习支持。配套发布预发布数据集S&I Corpus 2025,包含约315小时第二语言英语学习者口语数据,含整体评分;其中55小时有手动转录和语法错误标注。挑战设四项共享任务:自动语音识别(ASR)、口语语言评估(SLA)、口语语法纠错(SGEC)及纠错反馈(SGECF)。每项任务均设封闭赛道(限定模型与数据源)和开放赛道(允许使用公开资源)。参赛者可参与一项或多项任务。本文详述挑战设置、数据集及基准系统。
原文摘要 · Abstract (English)
This paper presents the "Speak & Improve Challenge 2025: Spoken Language Assessment and Feedback" -- a challenge associated with the ISCA SLaTE 2025 Workshop. The goal of the challenge is to advance research on spoken language assessment and feedback, with tasks associated with both the underlying technology and language learning feedback. Linked with the challenge, the Speak & Improve (S&I) Corpus 2025 is being pre-released, a dataset of L2 learner English data with holistic scores and language error annotation, collected from open (spontaneous) speaking tests on the Speak & Improve learning platform. The corpus consists of approximately 315 hours of audio data from second language English learners with holistic scores, and a 55-hour subset with manual transcriptions and error labels. The Challenge has four shared tasks: Automatic Speech Recognition (ASR), Spoken Language Assessment (SLA), Spoken Grammatical Error Correction (SGEC), and Spoken Grammatical Error Correction Feedback (SGECF). Each of these tasks has a closed track where a predetermined set of models and data sources are allowed to be used, and an open track where any public resource may be used. Challenge participants may do one or more of the tasks. This paper describes the challenge, the S&I Corpus 2025, and the baseline systems released for the Challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。